Remove the implicit hermes peer and the peer question from new connection setup. Preserve explicit peer settings and keep memory paths consistent with the captured client identity.
Add setup, configuration, request, recall, and session regression tests, plus upgrade guidance.
Follow-ups on the salvaged #97797 transport:
- Blocker 2 from the #97681 exact-head review claimed near-expiry refresh
silently mints against the target's CURRENT policy. On this head the
handler DOES refuse drift (_require_unchanged_execution_policy -> 403
room_reauthorization_required), but nothing pinned the handler-level
behavior: removing the drift check still passed the entire grants suite
(the check was only unit-tested in isolation). New HTTP-level regression
test drives /v1/room-members/grants/refresh with a drifted-policy grant
and requires the 403; sabotage-verified (check removed -> test fails).
- cancel() conflict resolution: keeps our race-retry routing loop from
#99099 with this layer's peer-stop acknowledgement body inside it
(peer receipt -> settle completion -> local interrupt escalation).
- docs: NAT one-way-reachability note in bot-mode.md — Desktop is a viewer,
not a relay; put room authority on the host everyone can reach (field
finding from /bin/bash on #97681).
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
ring entry removed from keyless_mcp, legacy backend set / credential
ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
nous_subscription surfaces. The tvly- redaction pattern stays --
legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
environment-variables, tools-reference, web-dashboard, provider
plugin dev guide.
Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
A delegate_task child that died (provider 404/400, timeout, crash)
previously vanished silently: the child's conversation loop returns
failed=True with the error summary in final_response, which the
classifier treated as usable output -> status 'completed'. And even
correctly-failed children only reached the parent MODEL — platforms
with tool_progress off (Telegram/Slack defaults) never showed the
human anything.
- delegate_tool: result.failed now forces status 'failed' (with the
error carried on the entry); new shared format_subagent_failure_line()
renders one clean human-readable line (traceback -> exception message,
length-capped); CLI tree + batch ✗ lines now include the reason.
- gateway TurnRunner.progress_callback: subagent.complete events with a
terminal failure status deliver that line via _deliver_platform_notice
BEFORE all progress-queue gates; tool_progress_callback is now always
attached (body gates each event class itself).
- tests: failed-flag classification regression + notice rendering suite.
- docs: Failure Visibility section in delegation docs.
Follow-ups on top of #98964's cherry-pick:
- PHOTON_READ_RECEIPTS env toggle (default true) so users can keep
messages at Delivered; declared in plugin.yaml optional_env
- adapter drops both 'read' and 'read_receipt' content types (alias
coverage from #91759 by @mooserini) + regression test
- docs: photon.md feature note + environment-variables.md row
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.
- The detailed session log is folded into the single summary request: the
lean prompt template gains a '## Detailed Session Log (oldest first)'
section carrying the digest prompt's HARD RULES (identifiers verbatim,
dense bullets, transcript-is-data). Output guidance grows by
_LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
summary budget — the old worst case (28 x 1,400 digest tokens) was spread
across many requests and mostly re-covered tool noise; a single dense
4K-token log inside one response preserves the load-bearing record while
staying well inside one aux response (the summary call still sends no hard
max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
whole region (_sample_summary_input: 8 proportionally spaced slices,
oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
anchored to the newest end) instead of head+tail truncated, so session-log
coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
_LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
_LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
attempt_summary_route_kwargs — no remaining callers; the single-use
summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
section lands in the summary; oversized regions sampled with elision
markers, never a second request; anchor index + recovery footer present).
Sabotage-verified: restoring a second call_llm makes the call-count test
fail. Docs and the compaction eval wording updated to stop claiming
per-chunk calls.
Fixes#96603.
Forward normalized custom-provider capabilities on the default gateway path so native compaction does not depend on session rehydration. Document the content trust boundary and cover both lookup and gateway resolution.
* refactor(skills): shipped-set slim — 15 skills to optional, github six-way merge, pdf absorbs OCR+nano-pdf, channel-gated teams pipeline
Maintainer-directed shipped-skills curation (skills index 1,900 -> ~1,400
tok/call on desktop; every session pays the index, so this is a per-call
diet on all installs):
- optional-skills moves (installable via skills hub, history preserved):
creative comfyui/ascii-art/excalidraw/pretext/sketch/touchdesigner-mcp;
ALL of mlops (huggingface-hub, llama-cpp, serving-llms-vllm,
weights-and-biases, evaluating-llms-harness — subcategory structure
kept); research-paper-writing (55 supporting files, 17.3K-tok load);
openhue; blogwatcher (first taught the cronjob monitor-field watch
pattern + web_extract instead of pre-cron manual workflows)
- DELETED session-librarian (Aug-12 'inspired by Perplexity Computer'
port, never maintainer-intended; session_search covers discovery)
- github: six skills (auth, issues, pr-workflow, issue-to-pr,
code-review, repo-management) merged into ONE software-development/
github skill — routing body + complete per-workflow references;
benbarclay authorship credited; codebase-inspection rides along;
discipline pins from test_github_issue_to_pr_skill.py preserved
against the reference body in the new test_github_skill.py
- pdf absorbs ocr-and-documents + nano-pdf as references/ + scripts
(extract_pymupdf, extract_marker converted to the argparse house
standard its contract test enforces)
- NEW session_platforms frontmatter gate (metadata.hermes): hides a
skill from the index on gateway channels it is not for; fail-open on
unknown platform; teams-meeting-pipeline gated to [teams, cron]
- blocked-page-recovery: research -> new web category; trigger-first
description ('Use when a fetch fails: 403/429, paywall, WAF, bot
wall.') so the model actually reaches for it on blocked fetches
- docs regenerated via generate-skill-docs.py (195 pages); related_skills
swept repo-wide; tests: 1672 passed (2 openclaw failures pre-existing
on clean main, Windows-local)
* chore: ignore .skills_prompt_snapshot.json (local index cache, accidentally committed)
Bundled by #83063 into the always-shipped set, but it serves only
multi-agent kanban campaigns — too niche for every install's skills
index (~20 tok/call for all users). Moved to the optional catalog
(installable via /skills search + hub, official source) and renamed so
the trigger is legible at a glance: 'merge-reconciler' read like a
generic git helper; 'agent-merge-conflict-arbiter' says who it is for.
- skills/autonomous-ai-agents/merge-reconciler -> optional-skills/autonomous-ai-agents/agent-merge-conflict-arbiter
- frontmatter name + description updated (description within the 60-char hardline its own contract test enforces)
- contract test moved/renamed, 9/9 green
- kanban docs (en + zh-Hans) repointed; zero merge-reconciler refs remain
Telegram's BotCommand menu is hard-capped (100/scope, ~4KB payload; Hermes
defaults to 60 slots), so most skill commands can never appear in the /
menu. Inline mode has no such cap: typing @botname <query> in any chat now
returns a live, searchable picker over EVERY core command, plugin command,
and installed skill — results computed per keystroke, paginated 50 at a
time. The Telegram analog of Discord's dynamic /skill autocomplete
(#18741).
- plugins/platforms/telegram/inline_picker.py: PTB-free catalog/rank/
pagination logic (unit-testable without python-telegram-bot). First
query token filters; the remainder is carried into the sent command as
its argument (@bot plan migrate auth → sends /plan migrate auth).
- adapter: InlineQueryHandler registration (inert until the bot owner
enables inline mode via BotFather /setinline) + _handle_inline_query
with the same auth path as inline-button callbacks — unauthorized users
get an empty list, so the installed-skill catalog is not leaked to
arbitrary users (inline queries arrive from any chat).
- Tap-to-send dispatches through the existing command path: the sent
message starts with /, which reaches the bot even under default privacy
mode. Zero new dispatch code.
- Docs: telegram.md inline-picker section incl. the one-time BotFather
/setinline setup.
Real-profile browsing (browser.use_real_profile) is meant to drive a COPY of
the user's profile headlessly in the background so they can keep working while
the agent acts on their behalf. Instead, on any host with a display it launched
the user's real browser binary HEADED, popping a window that grabbed focus on
every turn.
Root cause: the native launch (which bypasses agent-browser's own launcher to
avoid --use-mock-keychain dropping keychain-encrypted cookies) only added
--headless=new when Linux had no DISPLAY/WAYLAND_DISPLAY. On a normal desktop
the guard was false, so Chrome opened a visible window.
Fix: launch headless by default on every platform. Chrome's NEW headless mode
shares the profile's normal cookie store (unlike legacy --headless), and the
cookie drop we guard against comes from --use-mock-keychain, not from headless
— so real-profile auth still loads. Users who want to watch can opt in via the
existing browser.headed / AGENT_BROWSER_HEADED toggle (honored for real-profile
now, same as the rest of the browser stack); display-less hosts stay headless
regardless so the launch doesn't die at startup.
Live A/B on a real X seat: old argv mapped a Chrome window (focus steal),
--headless=new mapped zero windows while still exposing a working CDP port.
Updated the stale test that pinned 'never passes --headless' (a legacy-headless
premise) to positively assert the chrome launch is --headless=new by default.
The gateway's session-backed MCP OAuth flow (mcp.servers.oauth.start) binds
its browser-callback listener on the BACKEND machine's 127.0.0.1. When the
Desktop app connects to a remote backend (SSH/Tailscale), the user's browser
resolves that loopback to the user's machine, the redirect dies, and every
OAuth catalog server (ClickUp, Hospitable, ...) fails in-app with no working
path — the exact topology from the 'MCP Recurring erros' support thread.
Fix mirrors the Desktop's native gateway login (native-oauth-login.ts):
- gateway: mcp.servers.oauth.start accepts client_redirect_uri (loopback-only,
RFC 8252-style validation); when supplied no gateway listener is bound and
the OAuth redirect_uri pins to the client's listener.
- gateway: new mcp.servers.oauth.callback RPC relays the client-captured
code/state into the flow; state verification stays in
DashboardOAuthFlow.deliver_callback (constant-time compare, replay-safe).
- desktop: mcp-oauth-callback-ipc.ts hosts a one-shot 127.0.0.1 listener in
the main process (hermes:mcp-oauth:listen/wait/cancel via preload bridge).
- desktop: hermes-bots mcp-setup.tsx prefers the client listener for local
AND remote backends, falling back to the legacy gateway-listener flow on
older gateways (feature-detect via start rejection).
- docs: remote-host MCP OAuth section documents the automatic Desktop path.
Validation: 19 new gateway tests (validator allowlist, listener skip, relay
accept/reject/replay) — sabotage-verified; 5 new desktop tests against a real
ephemeral listener; E2E through the real session registry + flow bridge with
a stubbed provider probe; tsc electron+renderer builds clean.
platforms field added, description under the 60-char hardline, dangling
premium-webapp-ui related_skills ref dropped (user-profile skill, not in
the repo tree). Docs page + catalog row regenerated to match.
hermes skills install impeccable (and the docs-page install button) now
installs the impeccable frontend-design skill as an official optional-skills
entry. The local optional-skills/creative/impeccable/ dir is a catalog STUB:
its frontmatter declares metadata.hermes.upstream (repo + path), and
OptionalSkillSource.fetch() pulls the real 163-file bundle live from
pbakaus/impeccable:.hermes/skills/impeccable — the Hermes-native bundle
upstream maintains and verifies. Nothing vendored, never stale.
New mechanism (generic, not impeccable-specific):
- OptionalSkillSource._upstream_pointer(): parses/validates the upstream
pointer (owner/name repo, clean relative path, traversal rejected).
- _fetch_from_upstream(): delegates to GitHubSource.fetch(), relabels the
bundle official/<rel> at trust 'trusted' (curated endorsement, but
third-party content — dangerous scan verdicts still block).
- The live-repo fallback path redirects stubs the same way, so stale local
checkouts behave identically.
Three real gaps this surfaced, all fixed:
- GitHubSource.fetch() only downloaded SKILL.md plus paths linked from a
canonical support dir (references/, scripts/, ...). Impeccable keeps its
playbooks under reference/ (singular) and links scripts only from
reference files, so fetch shipped 1 of 163 files. fetch() now downloads
the full skill directory via the git tree (same approach as the
optional-skills live fetch), still rejecting symlinks/hidden/unsafe paths
and still failing on a missing SKILL.md-linked references/ path.
- The five env_exfil_* scanner patterns flagged loopback requests as
critical exfiltration: impeccable's live mode polls
http://localhost:PORT/status?token=TOKEN and scored two CRITICALs.
Scheme-anchored loopback exemption added; evil.com/?u=localhost decoys
still fire (10-case regex matrix in tests).
- unified_search() truncated to limit before ranking, so official catalog
entries got crowded out by skills.sh mirrors and bare-name installs
stalled on an ambiguity table. Results now stable-sort by trust rank
before the cut, and _resolve_short_name prefers a sole official exact
match over community mirrors.
Also fixes pre-existing test pollution: TestInstallPathSafety's fixture
monkeypatched the PEP 562 dynamic SKILLS_DIR, permanently shadowing dynamic
resolution and breaking the served_repo E2E tests in any combined run
(reproducible on main).
Validation: live E2E do_install("impeccable") against real GitHub —
resolves to official/creative/impeccable, verdict SAFE, 163 files on disk,
skill loads, /impeccable slash command registers. 128/128 targeted tests;
full-dir fetch test sabotage-verified. Docs: optional-skills catalog row,
generated skill page, sidebar.
The bundled plan skill's auto-generated slash command fell off the capped
Telegram/Discord command menus for most installs (skills are the only tier
trimmed at the platform caps, alphabetically — 'plan' sat past the cutoff at
index 57 of 82 bundled skills). Converting it to a first-class CommandDef
gives it a guaranteed core-tier menu slot on every platform.
- agent/plan_prompt.py: build_plan_prompt() — plan-mode rules + authoring
craft distilled from the retired skill; prompt-injection pattern like
/learn and /init (no engine, no model-tool footprint, cache-safe).
- CLI: _handle_plan_command mixin handler (pending-input injection).
- Gateway: /plan branch rewrites event.text and falls through (role
alternation preserved).
- TUI: command.dispatch branch ('plan' was already in
_PENDING_INPUT_COMMANDS).
- Removed skills/software-development/plan/ + docs pages (EN + zh-Hans),
catalog rows, sidebar entry, related_skills references.
- PROTECTED_BUILTIN_SKILLS is now empty (mechanism kept); dependent
curator/usage tests moved to monkeypatched sentinels.
Salvages #67292 by @webtecnica (credit: first /plan command submission,
issue #67264); reworked from inline planning prompt to the prompt-injection
pattern with workspace-saved plans. Closes#67264, closes#36821 (empty
/plan infers task from conversation context).
The grant_existing_profile key was removed in PR #98057; sweep the mode
table, opt-in section, runtime-lifecycle notes, config example, and CLI
reference that still documented it.
#95620 removed computer_use's cua_browser_* actions entirely
(browser_route.py, the action enum, and their dispatch/escalation hint) —
computer_use is desktop-only now, and page content goes through the
separate browser_navigate/browser_click/... toolset (or browser_exec
under the Browser Use CLI backend).
The skill doc never caught up: it still taught the model to call
cua_browser_state/cua_browser_prepare/etc. and to escalate to a "page"
rung that _enrich_escalation can no longer recommend. Following that
guidance fails schema validation on the first call.
- SKILL.md: replace the "Typed browser page rung" section (including the
existing_profile authorization walkthrough, which described a
model-facing action parameter that no longer exists anywhere in
schema.py/tool.py) with a short pointer to the current browser toolset;
drop "page" from the escalation.recommended union and the ladder step
that referenced it.
- website/docs/user-guide/features/computer-use.md: grant_existing_profile
and the rest of the permission-mode config are still real and current
(cua_backend.py still reads them for the runtime launch grant) — only
reworded the one sentence naming cua_browser_prepare as the mechanism
that consumes the grant, since driving a signed-in browser window now
goes through the standard capture/click/type actions instead.
- Regenerated the mirrored skill doc via
website/scripts/generate-skill-docs.py rather than hand-editing it, per
its own header. Kept the diff scoped to computer-use only.
Completes the #90953 salvage on post-#98237 main:
- New _merge_request_overrides helper defines the precedence contract:
explicit delegation.request_overrides merges OVER runtime/parent-derived
overrides — explicit top-level keys win; extra_body is deep-merged one
level so runtime extra_body keys survive unless redefined. Inputs are
copy.deepcopy'd so transport-side mutation can't leak into config or the
provider runtime cache.
- Direct base_url branch: explicit key now merges over the #98237
provider-alongside-base_url runtime overrides instead of being a separate
return shape; max_output_tokens preserved.
- Named-provider branch and parent-inherit branch now honor the key too, so
delegation.request_overrides never silently no-ops.
- _build_child_agent honors override_request_overrides whenever set
(previously only when override_provider was set), enabling the inherit
branch's merged value to reach the child.
- DEFAULT_CONFIG: delegation.request_overrides entry with comment.
- Tests: expanded tests/tools/test_delegate_request_overrides.py — deep-copy
proofs, explicit-over-runtime precedence on the provider-alongside-base_url
path, named-provider branch, inherit branch, and merge-helper unit tests.
- Docs: configuration.md delegation section + features/delegation.md document
the key, precedence, and example YAML (OpenRouter extra_body.provider.sort).
Extends PR #98250's classic-CLI status-bar upgrades to the Ink TUI:
- tui_gateway/server.py _get_usage() now emits cache_hit_pct,
avg_latency_s, avg_tps (reads the same per-call deque history from
agent/conversation_loop.py; keys omitted when no data — Codex
app-server has no latency, zero cache reads show no %)
- StatusRule renders the three read-outs as width-budgeted tail
segments (breakpoints 96/104/110 cols, lowest priority — they shed
first on narrow terminals)
- display.status_bar.fields (the SAME key the classic CLI honors)
filters TUI segments too: cache_hit, latency, tps, duration,
compressions, bg_tasks, bg_subagents, voice, battery, title,
context_pct, context_detail
- values ride the existing usage payload/ticker; constants between
events so the usage==last dedup keeps suppressing repaints
- 3 new server tests, 5 new TUI tests; full ui-tui suite 1727 green
Users reported no GUI switch for browser.use_real_profile — the only
desktop home was the generic Settings → Config editor, which nobody
found. The Browser toolset detail pane now renders a 'Use My Real
Browser Profile' ToggleRow above the backend/provider matrix.
- new BrowserRealProfilePanel: reads the shared profile-scoped config
record cache, optimistic write-through, rollback on failure
- saveHermesConfigRecord: capability-scoped PUT /api/config counterpart
of getHermesConfigRecord, so the Capabilities scope selector writes
the profile it points at (possibly another gateway)
- i18n: en/ja/zh/zh-hant keys (ar inherits en via defineLocale)
- docs: browser.md desktop pointer corrected to the real location
Live E2E on the built app over CDP: clicking the switch flipped
browser.use_real_profile true→false→true in the sandbox HERMES_HOME
config.yaml, GET reflected it, no layout glitches (screenshots in PR).
Profiles can now grant narrowly scoped tools to the background review runtime whitelist while unrelated tools remain denied. Document the configuration and cover it with a real-config regression test.
Agent: codex
- _terminate_real_profile_chrome(): directly-launched real browsers are ours
to reap (agent-browser only attaches); wired into the atexit emergency
cleanup and both launch-failure paths so orphaned Chrome processes can't
accumulate.
- Display-less Linux gate: append --headless=new (shares the profile's normal
cookie store, unlike legacy headless) so the direct-launch path doesn't
regress servers without DISPLAY/WAYLAND_DISPLAY.
- Register browser.real_profile_pin in config_defaults.py and document the
new launch model + pin in website/docs/user-guide/features/browser.md.
- Drop unused tempfile import from the cherry-picked commit.
Follow-up to the cherry-picked #41909/#92696 + #39760 + #97970 cluster:
- single field-key namespace (display.status_bar.fields) instead of the
second tui_statusbar_fields list; cache_hit/latency/tps/stash/battery/
title join the existing key set
- cache-hit % prefers the baseline-delta regime (resets on model switch
and compression) and hides on zero cache reads instead of alarming 0%
- latency/tps segments added to the styled fragment renderer too
- docs updated in website/docs/user-guide/configuration.md
- 7 new tests: rolling latency/t/s, NaN/negative guard, field filtering,
baseline resets, title badge gating
Allow users to control which fields appear in the interactive CLI status
bar via display.status_bar.fields in config.yaml.
Available fields: model, context_pct, context_detail, compressions,
bg_tasks, bg_processes, duration, prompt_elapsed, yolo, total_tokens.
When the list is empty (default), all fields are shown as before.
The field order is fixed (model always first); the config controls
visibility only. Narrow terminals (<76 cols) automatically drop
context_detail regardless of config.
total_tokens is opt-in only (not shown by default) to avoid width
overflow in the prompt_toolkit fragment renderer.
Closes#41909
Extends the real-profile machinery (PR #95620) to Brave Origin — Brave's
standalone paid build with a fully separate install identity:
- new canonical key 'brave-origin' in _CHROMIUM_BROWSERS
- Windows: BraveOHTML ProgId -> brave-origin; channel ProgIds BraveOBHTML/
BraveODHTML/BraveOSHTM fail closed (identifiers from brave-core
install_static)
- macOS: com.brave.Browser.origin bundle id (exact match); .beta/.dev/
.nightly channel bundles fail closed; /Applications/Brave Origin.app
- Linux: brave-origin.desktop matched BEFORE the bare 'brave' fragment
(substring scan would otherwise resolve an Origin default to stable
Brave and drive the wrong profile — #95549 wrong-principal invariant);
brave-origin-{beta,nightly,dev} fail closed
- profile dirs: BraveSoftware/Brave-Origin on all three OSes (per
brave-core kProductPathName + Homebrew cask zap paths)
- /browser connect launch tables: Brave Origin split into its OWN group
so a 'brave' executable lookup can never resolve to the Origin binary
- user-facing strings/docs/desktop tooltip updated
Tests: progid/bundle/desktop map params + data-dir resolution for all
three OSes; 125 passed in the three browser test files.
DEFAULT_AGENT_IDENTITY was rewritten in agent/prompt_builder.py (behavior
spec, exploration-thrift line deliberately removed) but the actual seed
written to disk on first run, hermes_cli/default_soul.py's
DEFAULT_SOUL_MD, was never updated. ensure_hermes_home() writes
DEFAULT_SOUL_MD into SOUL.md on every fresh install before the agent's
first turn, so virtually all real users end up as "SOUL.md users" seeded
with the pre-rewrite text -- including the exact "targeted and efficient
exploration" line the rewrite explicitly banned -- while the new
DEFAULT_AGENT_IDENTITY fallback essentially never serves the "fresh
install" audience its own PR body named as the target.
- DEFAULT_SOUL_MD now matches DEFAULT_AGENT_IDENTITY exactly.
- The pre-rewrite text is added to _LEGACY_TEMPLATE_SOULS so installs
already seeded with it self-heal via the existing upgrade-in-place
mechanism (same guarantee as the comment-only scaffold entries: the
string carries zero user intent, so it's safe to replace).
- Synced the other places install.sh's own comment says "MUST match
DEFAULT_SOUL_MD": scripts/install.sh, scripts/install.ps1,
docker/SOUL.md, and the docs/i18n pages that quote the fallback text
verbatim.
Live-reproduced on main: /handoff poll-waited a flat 60s for a TERMINAL
state, but the gateway's dispatch is a full synthetic agent turn (whole
transcript replay + delivery) that routinely exceeds 60s on long sessions.
The CLI then printed "Timed out waiting for the gateway. Is `hermes
gateway` running?" (false diagnosis), called fail_handoff() on the RUNNING
row (stomping the gateway's claim), and promised "Your CLI session is
intact" after switch_session had already re-pointed the session. The
watcher later overwrote failed -> completed: split-brain.
- hermes_state.fail_handoff gains only_states CAS; waiters can only fail
rows still pending. Owner (gateway watcher) keeps the unconditional form.
- CLI wait loop is two-phase: 60s for the CLAIM (pending) — a timeout
there really does mean no gateway — then up to 15 min for the claimed
dispatch with 30s heartbeats; a running row is never failed by the CLI.
- Desktop handoff.fail RPC now CAS-fails pending rows only; a running row
returns {failed: false, state: running} instead of stomping the claim.
Repro (real _handoff_watcher, real state.db, CLI as separate process,
75s dispatch): before — CLI timeout @60s + false message + row stomped;
after — pending->running@5s->completed@80s, clean CLI exit.
Sweeper review, all three points:
- Placement: no new plugins/model-providers/ directory. The Token Plan
profiles register from the existing alibaba plugin module — one module
per vendor, matching how the kimi module carries both of its endpoint
variants. Token Plan is the same vendor/service (Model Studio), same
OpenAI-compatible protocol, its own key + endpoints; splitting to a
standalone repo remains a 5-minute change if maintainers prefer.
- Runtime coverage: TestRuntimeAlibabaRegionalAndTokenPlan exercises
resolve_runtime_provider() for all four variants — provider, api_key,
api_mode, base_url — alongside the existing zai/minimax/kilocode
runtime regressions.
- Docs: providers.md, environment-variables.md, cli-commands.md updated
with the bundled variants and their env keys/base-url overrides.
Flip the salvaged --start-now behavior (PR #97958) into the unconditional
default: /loop's first iteration is due the moment the loop is set, then
recurs on the normal cadence. The flag is dropped — it was never released,
so there is nothing to deprecate.
- LoopManager.set(): next_due_at = now for both cadence modes
- drop --start-now parsing, the persisted LoopState.start_now field, and
the flag from help text; confirmation now always says the first wakeup
fires now
- tests updated to pin the new default (incl. the TUI not-due test, which
now has to push next_due_at out explicitly)
- docs: quick-start and command table describe the immediate first run
/loop [interval] <prompt> currently schedules the first wakeup one full
interval after the command runs (next_due_at = now + interval). When the
user just told Hermes what to check, waiting the whole interval before
any output feels like the command was ignored.
Add an opt-in --start-now flag that keeps Claude Code parity as the
default but lets the user run the first iteration immediately, then
continue on the cadence:
/loop 1h check the deploy status # first run in 1h (unchanged)
/loop 1h --start-now check the deploy # first run now, then hourly
- parse_loop_args(): parse and strip --start-now (leading or trailing)
- LoopState: new persisted start_now field (default False, survives
serialization round-trip and old rows missing the field)
- LoopManager.set(): next_due_at = now when start_now, for both fixed
interval and self-paced modes
- dispatch_loop_command(): wire start_now through, update help text, and
report "First wakeup fires now" in the confirmation
- website/docs: document the flag in the /loop guide
- tests: parse (trailing/leading/absent/self-paced/combo/prompt-word),
tick lifecycle (due immediately vs after interval), serde round-trip,
and dispatch-level confirmation
Router shipped a minimal /v1/chat/completions compatibility surface
(translated onto Responses) after this PR was written, so the
'does not exist and 404s' wording is stale. Responses remains the
native wire — per-model reasoning-effort validation, reasoning
summaries, and prompt caching live there — so the api.router.com
host mandate is unchanged; only the comments and docs are updated.
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:
- plugins/model-providers/router/: RouterProfile plugin —
api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
runtime_provider._detect_api_mode_for_url: api.router.com ->
codex_responses. The host is Responses-only — POST /v1/chat/completions
does not exist and 404s — so this is a genuine host mandate (exact
hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
hook (tri-state: None=defer, ()=model takes no reasoning params,
tuple=clamp set). Router validates reasoning.effort per model and
returns HTTP 400 invalid-argument on levels outside the model's
published vocabulary, and 400 unsupported_parameter when a
non-reasoning model receives any reasoning field (both verified live).
The profile answers from a cached copy of the catalog's
router.capabilities.reasoning block: cache-only on the hot path,
seeded for free by fetch_models(), disk-mirrored across processes
(/cache/router_catalog.json), background-warmed when cold
— same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
the generic effort-clamp branch (xai/actual/github branches untouched;
profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
rejection, profile registration + auth auto-registry wiring, catalog
parsing, and transport clamp/suppression/fallback paths.
Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.
/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.
Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).