Live-reproduced on main: /handoff poll-waited a flat 60s for a TERMINAL
state, but the gateway's dispatch is a full synthetic agent turn (whole
transcript replay + delivery) that routinely exceeds 60s on long sessions.
The CLI then printed "Timed out waiting for the gateway. Is `hermes
gateway` running?" (false diagnosis), called fail_handoff() on the RUNNING
row (stomping the gateway's claim), and promised "Your CLI session is
intact" after switch_session had already re-pointed the session. The
watcher later overwrote failed -> completed: split-brain.
- hermes_state.fail_handoff gains only_states CAS; waiters can only fail
rows still pending. Owner (gateway watcher) keeps the unconditional form.
- CLI wait loop is two-phase: 60s for the CLAIM (pending) — a timeout
there really does mean no gateway — then up to 15 min for the claimed
dispatch with 30s heartbeats; a running row is never failed by the CLI.
- Desktop handoff.fail RPC now CAS-fails pending rows only; a running row
returns {failed: false, state: running} instead of stomping the claim.
Repro (real _handoff_watcher, real state.db, CLI as separate process,
75s dispatch): before — CLI timeout @60s + false message + row stomped;
after — pending->running@5s->completed@80s, clean CLI exit.
Sweeper review, all three points:
- Placement: no new plugins/model-providers/ directory. The Token Plan
profiles register from the existing alibaba plugin module — one module
per vendor, matching how the kimi module carries both of its endpoint
variants. Token Plan is the same vendor/service (Model Studio), same
OpenAI-compatible protocol, its own key + endpoints; splitting to a
standalone repo remains a 5-minute change if maintainers prefer.
- Runtime coverage: TestRuntimeAlibabaRegionalAndTokenPlan exercises
resolve_runtime_provider() for all four variants — provider, api_key,
api_mode, base_url — alongside the existing zai/minimax/kilocode
runtime regressions.
- Docs: providers.md, environment-variables.md, cli-commands.md updated
with the bundled variants and their env keys/base-url overrides.
Flip the salvaged --start-now behavior (PR #97958) into the unconditional
default: /loop's first iteration is due the moment the loop is set, then
recurs on the normal cadence. The flag is dropped — it was never released,
so there is nothing to deprecate.
- LoopManager.set(): next_due_at = now for both cadence modes
- drop --start-now parsing, the persisted LoopState.start_now field, and
the flag from help text; confirmation now always says the first wakeup
fires now
- tests updated to pin the new default (incl. the TUI not-due test, which
now has to push next_due_at out explicitly)
- docs: quick-start and command table describe the immediate first run
/loop [interval] <prompt> currently schedules the first wakeup one full
interval after the command runs (next_due_at = now + interval). When the
user just told Hermes what to check, waiting the whole interval before
any output feels like the command was ignored.
Add an opt-in --start-now flag that keeps Claude Code parity as the
default but lets the user run the first iteration immediately, then
continue on the cadence:
/loop 1h check the deploy status # first run in 1h (unchanged)
/loop 1h --start-now check the deploy # first run now, then hourly
- parse_loop_args(): parse and strip --start-now (leading or trailing)
- LoopState: new persisted start_now field (default False, survives
serialization round-trip and old rows missing the field)
- LoopManager.set(): next_due_at = now when start_now, for both fixed
interval and self-paced modes
- dispatch_loop_command(): wire start_now through, update help text, and
report "First wakeup fires now" in the confirmation
- website/docs: document the flag in the /loop guide
- tests: parse (trailing/leading/absent/self-paced/combo/prompt-word),
tick lifecycle (due immediately vs after interval), serde round-trip,
and dispatch-level confirmation
Router shipped a minimal /v1/chat/completions compatibility surface
(translated onto Responses) after this PR was written, so the
'does not exist and 404s' wording is stale. Responses remains the
native wire — per-model reasoning-effort validation, reasoning
summaries, and prompt caching live there — so the api.router.com
host mandate is unchanged; only the comments and docs are updated.
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:
- plugins/model-providers/router/: RouterProfile plugin —
api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
runtime_provider._detect_api_mode_for_url: api.router.com ->
codex_responses. The host is Responses-only — POST /v1/chat/completions
does not exist and 404s — so this is a genuine host mandate (exact
hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
hook (tri-state: None=defer, ()=model takes no reasoning params,
tuple=clamp set). Router validates reasoning.effort per model and
returns HTTP 400 invalid-argument on levels outside the model's
published vocabulary, and 400 unsupported_parameter when a
non-reasoning model receives any reasoning field (both verified live).
The profile answers from a cached copy of the catalog's
router.capabilities.reasoning block: cache-only on the hot path,
seeded for free by fetch_models(), disk-mirrored across processes
(/cache/router_catalog.json), background-warmed when cold
— same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
the generic effort-clamp branch (xai/actual/github branches untouched;
profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
rejection, profile registration + auth auto-registry wiring, catalog
parsing, and transport clamp/suppression/fallback paths.
Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.
/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.
Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
The todo tool now supports hierarchical task lists: an item's optional
'parent' field points at another item's id, making it a subtask.
- tools/todo_tool.py: parent validated (self-ref dropped), dangling refs
and cycles sanitized; merge mode can set/clear parent; post-compression
injection renders the tree indented and keeps a finished parent visible
while any descendant is still active; the in-progress reorder pass is
skipped for nested lists (a flat move would tear subtasks from parents).
- Schema cost: ~45 tokens added to the cached tool schema (one string
property + one behavior sentence).
- acp_adapter/tools.py: todo result markdown indents by parent depth.
- Desktop: TodoItem carries parent; todoTree() DFS helper; composer
status stack renders subtask rows indented (depth-capped), stabilizer
compares depth.
- Docs: tools-reference todo entry mentions nesting.
Hydration/replay paths (gateway fresh-agent, API-server history) work
unchanged: parent rides inside the same todos array.
Reverts the plan-interrogation rename from #97831 per maintainer decision —
the skill keeps its original grill-me name. The content upgrade (design-tree
frontier-rounds interview mechanic from mattpocock/skills' grilling) stays.
Docs page, catalog row, and sidebar entry renamed back.
Ports the MIT-licensed 'to-questionnaire' skill from mattpocock/skills as
decision-questionnaire (optional-skills/productivity). Interviews the user
about the send only (recipient + needed outcomes), then drafts a
most-important-first discovery questionnaire with answer stubs to
decision-questionnaire-<slug>.md. Includes staging tests for frontmatter
standards, template sections, and de-upstreaming.
Ports the MIT-licensed 'wizard' skill from mattpocock/skills as
setup-wizard-generator (optional-skills/devops). Generates an interactive
bash wizard that walks a human through manual procedures: opens dashboard
URLs, captures values (hidden entry for secrets), writes .env / GitHub
secrets idempotently, and confirms each stage. Vendors upstream's
template.sh library verbatim (bash -n verified) plus a staging test suite
covering frontmatter, template integrity, and de-upstreaming.
Renames the optional grill-me skill to the more descriptive plan-interrogation
and upgrades it with the design-tree frontier-rounds mechanic from
mattpocock/skills' MIT-licensed 'grilling' skill: batch all currently-askable
questions per round in dependency order, agent finds facts itself, decisions
stay with the user. Docs page, catalog row, and sidebar entry renamed.
- Widen opencode_zen_free_runtime healing to the union of the static floor,
the in-process live memo, and the SWR disk cache — a newly-live free model
now heals opencode-go/zen selections without a release (sibling site the
original PR missed).
- Memoize _fetch_opencode_free_models() in-process (5 min, negative caching
included) so direct provider_model_ids() validation callers don't each
block on a network round-trip or timeout.
- Drop delisted x-preview-f-free from the offline floor and setup.py sample
list (offline fallback must not offer a model that 401s); add the newly
live deepseek-v4-flash-free / mimo-v2.5-free to setup.py.
- Update stale test fixtures to a live exemplar; add regression tests for
memoization, negative caching, and union healing; docs note in providers.md.
The quote in x-azed-tokyo-trip names no agent, and the linked post
credits the itinerary to MyClaw, mentioning Hermes only as another agent
that account manages. The collage promises stories about how people use
Hermes, so an entry that cannot be attributed to Hermes does not belong
in it.
Removed rather than reworded: the source does not support a
Hermes-attributed version of this story. 64 new entries remain.
Adds 65 user stories gathered from Reddit and X, prepended to the front of
the list so the newest entries surface first.
Every quote was re-fetched from its live source and programmatically verified
to be a verbatim substring of the source text. No existing entries are
modified, reordered, or removed - this is a pure prepend.
45 Reddit, 20 X. All categories and sources already exist in the file; no
schema change.
Per the 'when in doubt, optional' rule — plan-interview is an
on-request capability, not a weekly daily-driver for most users.
Install via: hermes skills install official/software-development/grill-me
Session hygiene auto-compression runs inline on the incoming-message path
and awaits the summary worker with a progress-aware inactivity budget
(hygiene_timeout_seconds) that extends up to hygiene_total_ceiling_seconds
(default 600s). A summary model that keeps streaming tokens keeps resetting
the inactivity slice, so the wait can stretch toward the ceiling while zero
bytes reach the user — chat transports (Telegram ~30s idle-timeout) drop the
connection and the turn appears frozen, even though the gateway is healthy.
Add hygiene_max_turn_hold_seconds (default 10), a turn-hold budget that caps
the wall-clock the incoming message waits on hygiene compression. The wait
slice is additionally capped at the remaining budget so the budget is
re-evaluated even when the worker keeps the inactivity slice large. On
exceeding the budget the gateway abandons the inline wait and proceeds on
the uncompressed transcript via the existing timeout path, which revokes the
worker's commit admission (CompressionCommitFence) and defers cleanup — so a
stale compression finishing later can never overwrite the turns appended
after the wait was abandoned.
Well under the typical transport idle-timeout, this guarantees the message
is answered promptly while the detached compression completes in the
background. Configurable via compression.hygiene_max_turn_hold_seconds.
Adds a regression test: a worker that streams progress continuously (so the
inactivity slice never fires) must be abandoned once it exceeds the
turn-hold budget, the turn proceeds uncompressed, and the stale commit is
fenced (no session mutation, role alternation intact).
The locked dependency tree now carries @babel/* 8.x, which requires
node ^22.18.0 || >=24.11.0. Our engines.node arm said ^24.0.0 and the
installer gates (node_satisfies_build / Test-NodeVersionOk) accepted any
Node 24 — so a system Node 24.0–24.10 cleared every gate we own and then
failed 'npm install' with EBADENGINE under engine-strict=true.
- Raise the 24 arm to ^24.11.0 in root + desktop package.json and the
package-lock.json mirrors
- Tighten node_satisfies_build (install.sh) and Test-NodeVersionOk
(install.ps1) to 24.11+; update user-facing wording
- Add invariant tests: every engines.node arm floor must satisfy every
locked dependency's engines.node, and the installer gates must encode
the same floors as the manifest — so the next babel-style floor bump
turns into a CI red instead of a user install outage
- docs: correct stale 'Node.js v22' provisioning claim
Follow-up on top of @rahlquist's terminal.temp_dir knob (#97182): the
default itself now avoids RAM-backed tmpfs. Resolution order on the
local backend: terminal.temp_dir > TMPDIR/TMP/TEMP > HERMES_HOME/cache/
terminal (managed, pruned) > /tmp fallback. Pruning: hourly via gateway
housekeeping + once-per-process best-effort sweep; hermes_bg_* triplets
are aged as a group so a live server's fresh .log protects its .pid.
On a real TTY, `hermes chat -q "…"` (and `--tui -q`) now starts a normal
interactive session with the prompt submitted literally as the first turn —
no slash-command routing, no '!' shell dispatch, no $(...) interpolation,
no file-drop rewriting — matching how other coding agents handle seeded
launches (Omarchy prompted agent terminals, basecamp/omarchy#8705).
Legacy answer-and-exit is preserved everywhere automation depends on it:
- new `hermes chat --oneshot` flag (distinct dest from top-level -z)
- -Q/--quiet machine-readable contract
- any non-TTY stdio (kanban workers, cron, pipes, A2A)
- top-level `hermes -z` unchanged
CLI: seeded prompt rides a _SeededQueryMessage sentinel through
process_loop, which skips the slash/!/file-drop dispatchers for that one
message. TUI: STARTUP_QUERY submits via a new literal path (submitLiteral)
that bypasses dispatchSubmission and the input.detect_drop rewrite.
Live on both providers (verified 2026-08-28 against openrouter.ai/api/v1/models
and inference-api.nousresearch.com/v1/models) but absent from both curated
picker lists. Adds the entry directly below qwen3.8-max per newest-first
family ordering, an explicit 1M DEFAULT_CONTEXT_LENGTHS entry (new family
slug would otherwise fall through to the generic qwen 131072 catch-all —
same class as #69881), and regenerates model-catalog.json.
Scoped rollout: only the named providers touched. Pricing snapshot skipped
(both routes bill via official_models_api live pricing). Reasoning floor
already fires via the qwen3 prefix entry (180s, verified).
Made opt-in when it fired a tip 45 seconds into a launch and another
every six minutes, which is a cadence that owes you a choice. The pacing
has since become a settling delay of five to ten minutes per launch and
a six-hour cooldown persisted across them — roughly a tip a day, weeks
to walk the catalog. At that weight the switch has nothing left to
protect anyone from, and a discovery feature nobody meets is one nobody
has. The switch stays for whoever still wants it off.
A first tip 45 seconds in and one every six minutes after walks the
whole catalog in an hour, which is the cadence of a notification rather
than a nicety. Games get this right by being almost absent: a tip while
you settle in, then nothing for the rest of the day.
Two clocks now have to agree. A per-launch settling delay of five to ten
minutes means opening the app is never met with a bubble, and a six-hour
cooldown persisted across launches means quitting and reopening isn't a
way to farm them — the old schedule lived in the effect and re-armed on
every mount. Flipping the switch on skips the settling delay and offers
immediately, since that clock guards a launch you came into with a
purpose, not a deliberate opt-in.
An agent tip starts the cooldown too: whoever just pointed at something,
the user has had their one interruption for a while.
The two halves of tips were behind one switch, which meant the app
volunteering commentary at idle shipped on by default. Split them along
the line that matters: the rotation talks unprompted, so it now waits to
be asked for, while an agent tip stays ungated like the tour it mirrors
— Hermes raises one mid-conversation, in answer to something the user
said.
Drops the tool's config gate along with the config key it read. The
renderer mirrored that key with config.set, which has no branch for it
and answered "unknown config key" into a swallowed catch, so the opt-out
never reached the backend in the first place.
The quiet sibling of `tour`, in the same `desktop_ui` toolset and reading the
same `tour(action='targets')` discovery call: one bubble with an arrow, for a
sentence that would be clearer with a finger on the thing it's about. Dimming
the whole app to say "the model name is a button" is the wrong weight.
Fire-and-forget rather than a round-trip, because a tip is not a question and
blocking the turn on one would stall the reply it belongs to. The renderer
enforces the user's opt-out itself, so a stale config read can never put a
bubble on a screen that asked for none.
A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.
The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.
- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
api_server/msgraph_webhook wire before their dispatch tables freeze;
the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
calling the hook.
Mirrors the Slack precedent (register_slack_action_handler): plugins queue
a factory at register() time; the Telegram adapter invokes each factory
with (application, adapter) at connect() time, before the core handlers
register, so pattern-scoped plugin handlers take precedence for their own
updates while everything else falls through unchanged. Factories are
isolated — a raising plugin cannot prevent Telegram from connecting.
Unblocks standalone plugins that need PTB update types the core adapter
doesn't route (Telegram Business API secretary bots, custom callback
prefixes, chat-member events) without touching core files.
- The direct adapter.send() confirmation path in /approve and /deny is only
needed on native-streaming platforms (WeCom) where the reply stream is
already finalized; other platforms keep the return-text contract (fixes
4 approve/deny regression tests, guarded with 'is not True' against
MagicMock auto-attributes).
- Remove a stray [DEBUG] logger.info left in _deliver_media_from_response.
- website/docs wecom.md: replace the 'does not stream' notes with the native
msgtype:stream behavior and document the stream keepalive extra keys.
Per the 'when in doubt, optional' rule — site publishing is an
on-request capability, not a weekly daily-driver for most users.
Joins cloudflare-temporary-deploy/page-agent under
optional-skills/web-development (existing category, existing
DESCRIPTION.md kept; the new bundled category dir is dropped).
Install via: hermes skills install official/web-development/publish-site
Live /api/v1/models probe (2026-08-27) confirms the id is gone from the
catalog, so the curated picker entry was a dead pick. Manifest
regenerated. No provider-agnostic metadata existed for the slug.
Delist credit: @orouge97 flagged this in PR #80036.
Allowlist hot-path hooks for abandon-on-timeout, keep subagent_stop on the caller thread, suppress re-fires of hung callbacks, and block tools when pre_tool_call times out.
- move minimax/minimax-m3:free into the Free tier section (house
convention: :free SKUs group together, matching glm-5.2:free and the
nemotron :free entries) and regenerate model-catalog.json
- add Inkling family context length (1,048,576 — OpenRouter live
metadata, 2026-08-27) to DEFAULT_CONTEXT_LENGTHS; new family slug
otherwise fell through to no entry
- add Inkling to the reasoning stale-timeout floor table (300s tier,
same as Grok reasoning / Ox Alpha; OpenRouter marks the family as
reasoning-capable)
- widen the floor matcher's right-anchor separator class to include
':' so OpenRouter SKU suffixes (:free/:batch/:nitro) inherit the
family floor — inkling:free previously missed the inkling entry
- regression tests for the inkling floor + ':' separator
The OpenRouter model picker builds its list from a curated set of model
IDs, then filters against OpenRouter's live catalog. minimax/minimax-m3:free
exists on OpenRouter (free tier, 1M context, tool-calling) but was missing
from both the in-repo fallback list and the remote catalog manifest.
Add it to OPENROUTER_MODELS and website/static/api/model-catalog.json so
the free variant surfaces in the picker alongside the paid one.