Follow-up for salvaged PR #28253: current main's generic profile fetch
passes base_url= to fetch_models and merges curated fallback_models
first (658ac1d86 / #46309), so the mocked-signature and exact-equality
assertions from the PR's era no longer match the contract.
Review findings on salvaged #93548: _warm_efforts_async now returns
early under PYTEST_CURRENT_TEST (matching the canonical OpenRouter caps
warmer) so a test that forgets to monkeypatch it can't fire live HTTP
when RAMP_ROUTER_API_KEY is set; the codex transport's fail-open
except in _profile_declared_efforts logs at debug instead of silently
swallowing profile-hook bugs.
Router shipped a minimal /v1/chat/completions compatibility surface
(translated onto Responses) after this PR was written, so the
'does not exist and 404s' wording is stale. Responses remains the
native wire — per-model reasoning-effort validation, reasoning
summaries, and prompt caching live there — so the api.router.com
host mandate is unchanged; only the comments and docs are updated.
Addresses the automated review on this PR:
- _profile_declared_efforts falls back from provider name to the
endpoint's host (via model_metadata's URL->provider map), so a named
custom provider pointed at api.router.com — which the host mandate
already routes onto this transport — gets the catalog clamp instead
of the default vocabulary and a Router 400.
- _parse_efforts validates catalog levels against EFFORT_LADDER at
ingest, logging and dropping unrecognized tiers; a model whose whole
vocabulary is unrecognized stays out of the map (transport defaults)
instead of passing the requested effort through unclamped.
- fetch_models dedupes ids while preserving Router's deliberate listing
order.
- plugin.yaml credits the human contributor per repo convention.
Router attributes coding-agent clients by User-Agent prefix (the way it
already recognizes OpenCode's versioned UA), and its WAF rejects
default/blank client UAs. Mirror the xai profile: declare
User-Agent: Hermes-Agent/<version> in default_headers, which
agent_init's profile-headers fallback applies at client construction.
Re-verified live: one-shot chat through api.router.com still works.
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:
- plugins/model-providers/router/: RouterProfile plugin —
api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
runtime_provider._detect_api_mode_for_url: api.router.com ->
codex_responses. The host is Responses-only — POST /v1/chat/completions
does not exist and 404s — so this is a genuine host mandate (exact
hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
hook (tri-state: None=defer, ()=model takes no reasoning params,
tuple=clamp set). Router validates reasoning.effort per model and
returns HTTP 400 invalid-argument on levels outside the model's
published vocabulary, and 400 unsupported_parameter when a
non-reasoning model receives any reasoning field (both verified live).
The profile answers from a cached copy of the catalog's
router.capabilities.reasoning block: cache-only on the hot path,
seeded for free by fetch_models(), disk-mirrored across processes
(/cache/router_catalog.json), background-warmed when cold
— same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
the generic effort-clamp branch (xai/actual/github branches untouched;
profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
rejection, profile registration + auth auto-registry wiring, catalog
parsing, and transport clamp/suppression/fallback paths.
Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.
Review finding (quality pass): on a single-profile gateway,
_handle_busy_command set _busy_input_mode but left _busy_text_mode
stale, so the adapter refresh a line later re-read the old value —
/busy queue persisted to config but live text messages kept
interrupting until restart. The profile path already re-derives both
from the fresh config; the non-profile path now does the same via
_load_busy_text_mode() (busy_input_mode is the source of truth,
run.py:9877). Regression assertion added to test_set_mode_persists;
verified red without the production fix.
test_update_is_known_command grepped _handle_message's source for the
literal '"update"' — a banned source-reading test (AGENTS.md), broken
by the if-chain -> _gateway_plain_command_handlers() refactor. Assert
the actual dispatch contract instead: the shared handler table maps
'update' to _handle_update_command.
Teknium review items:
1. Parse with event.get_command_args() instead of raw event.text
(matching _handle_fast_command pattern at line 2842)
2. Add mocked persistence tests for queue/steer/interrupt setter
success, save-failure, and exception branches (5 new tests)
Removes cli_only=True from /busy CommandDef and adds gateway
handler with subcommand dispatch (status/queue/steer/interrupt).
Applied on top of latest upstream/main while preserving original
commit intent from PR #18366.
Also adds smoke tests for the gateway /busy command handler.
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.
/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.
Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
The todo tool now supports hierarchical task lists: an item's optional
'parent' field points at another item's id, making it a subtask.
- tools/todo_tool.py: parent validated (self-ref dropped), dangling refs
and cycles sanitized; merge mode can set/clear parent; post-compression
injection renders the tree indented and keeps a finished parent visible
while any descendant is still active; the in-progress reorder pass is
skipped for nested lists (a flat move would tear subtasks from parents).
- Schema cost: ~45 tokens added to the cached tool schema (one string
property + one behavior sentence).
- acp_adapter/tools.py: todo result markdown indents by parent depth.
- Desktop: TodoItem carries parent; todoTree() DFS helper; composer
status stack renders subtask rows indented (depth-capped), stabilizer
compares depth.
- Docs: tools-reference todo entry mentions nesting.
Hydration/replay paths (gateway fresh-agent, API-server history) work
unchanged: parent rides inside the same todos array.
Folded from #97748 (the competing fix by @lEWFkRAD): when a rollback
restore fails, the error note now points the operator at the surviving
snapshot directory instead of leaving them to find it in tempdir.
Rollback removed the live skill directory before restoring its
snapshot. When copytree then failed (disk full, locked file, path too
long on Windows) the except only added a note, and the finally deleted
the snapshot directory too, so nothing survived: the skill was gone
with a success-shaped error payload.
The broken state is now renamed aside first and deleted only after the
snapshot is restored. If the restore still fails, the broken state is
renamed back, so the worst outcome is the half applied batch instead of
no skill at all. When rollback reports any failure the snapshots are
kept on disk and their location is logged, instead of being deleted by
the finally.
Follow-up to #97692, same batch executor.
* test(system_prompt): cover session-start anchoring + fix Windows-portable expectation
- Add TestSessionStartLike unit tests for _session_start_like(): session-id
embedded timestamp, session_start fallback, now fallback, non-matching id.
- Add a build-level regression: a session started Jan 1 must still render
'Conversation started: Thursday, January 01' when the prompt is rebuilt
on Jan 2 (the rebuild-drift bug).
- test_coding_prompt_preserves_legacy_workspace_order hardcoded '/hermes'
while production renders str(Path('/hermes')) — backslash on Windows made
the suite fail on Windows (CI runs Linux, so it was never caught). Build
the expectation via str(Path()) to match production on every platform.
* fix(system_prompt): anchor 'Conversation started' to the real session start
The timestamp line stamped hermes_time.now() at system-prompt build time.
The prompt is rebuilt on compression, fresh-agent gateway turns, and
resume-without-stored-prompt, so the date silently advanced to whatever
day the prompt was last rebuilt — a chat that started on Wednesday read
as 'Conversation started: Thursday' after a Thursday-morning resume,
contradicting the fresh per-turn time hint.
Resolve the true start via _session_start_like(): the timestamp embedded
in the session id (YYYYMMDD_HHMMSS_..., immutable for the session life)
-> agent.session_start -> now() only as last resort. Box-local stamps are
attached to the box's local zone then converted to the rendered zone so
the date is consistent with the per-turn clock. The line stays date-only
and is now byte-stable for the whole session (never moves on rebuild),
preserving prefix-cache KV. The zone suffix and _bot_chat_timeless_prompt
behaviour are untouched.
* feat(system_prompt): two-line conversation clock — anchored start (salvaged #96224, credit @bobaba76) + as-of-last-rebuild date for multi-day sessions
---------
Co-authored-by: bobaba76 <79245850+bobaba76@users.noreply.github.com>
Reverts the plan-interrogation rename from #97831 per maintainer decision —
the skill keeps its original grill-me name. The content upgrade (design-tree
frontier-rounds interview mechanic from mattpocock/skills' grilling) stays.
Docs page, catalog row, and sidebar entry renamed back.
Ports the MIT-licensed 'to-questionnaire' skill from mattpocock/skills as
decision-questionnaire (optional-skills/productivity). Interviews the user
about the send only (recipient + needed outcomes), then drafts a
most-important-first discovery questionnaire with answer stubs to
decision-questionnaire-<slug>.md. Includes staging tests for frontmatter
standards, template sections, and de-upstreaming.
Ports the MIT-licensed 'wizard' skill from mattpocock/skills as
setup-wizard-generator (optional-skills/devops). Generates an interactive
bash wizard that walks a human through manual procedures: opens dashboard
URLs, captures values (hidden entry for secrets), writes .env / GitHub
secrets idempotently, and confirms each stage. Vendors upstream's
template.sh library verbatim (bash -n verified) plus a staging test suite
covering frontmatter, template integrity, and de-upstreaming.
Renames the optional grill-me skill to the more descriptive plan-interrogation
and upgrades it with the design-tree frontier-rounds mechanic from
mattpocock/skills' MIT-licensed 'grilling' skill: batch all currently-askable
questions per round in dependency order, agent finds facts itself, decisions
stay with the user. Docs page, catalog row, and sidebar entry renamed.
- Widen opencode_zen_free_runtime healing to the union of the static floor,
the in-process live memo, and the SWR disk cache — a newly-live free model
now heals opencode-go/zen selections without a release (sibling site the
original PR missed).
- Memoize _fetch_opencode_free_models() in-process (5 min, negative caching
included) so direct provider_model_ids() validation callers don't each
block on a network round-trip or timeout.
- Drop delisted x-preview-f-free from the offline floor and setup.py sample
list (offline fallback must not offer a model that 401s); add the newly
live deepseek-v4-flash-free / mimo-v2.5-free to setup.py.
- Update stale test fixtures to a live exemplar; add regression tests for
memoization, negative caching, and union healing; docs note in providers.md.
opencode-free (keyless) models were served exclusively from a hardcoded
in-repo snapshot (_PROVIDER_MODELS["opencode-free"]). The SWR disk cache
only revalidated AUTHED providers — its entries were keyed by a credential
fingerprint, which keyless providers have none of — so the catalog never
refreshed against GET /zen/v1/models. When the relay delisted a free model
(e.g. x-preview-f-free, 2026-08-26) the picker kept offering it and
selecting it 401'd: "Model x-preview-f-free is not supported".
Now provider_model_ids("opencode-free") fetches the live /zen/v1/models
catalog anonymously, filters it to the anonymous-servable free tier
(excluding KEYED suffix-fakes like Go's ox-alpha-free), and falls back to
the curated static floor only when the live fetch fails or is empty. The
keyless provider gets a stable disk-cache fingerprint so the picker's SWR
path serves stale immediately while refreshing off-thread — the same
behavior authed providers already get.
Regression tests prove the fix: the delisted/newly-live model assertions
fail when the live-fetch wiring is reverted.
Closes#95914
The quote in x-azed-tokyo-trip names no agent, and the linked post
credits the itinerary to MyClaw, mentioning Hermes only as another agent
that account manages. The collage promises stories about how people use
Hermes, so an entry that cannot be attributed to Hermes does not belong
in it.
Removed rather than reworded: the source does not support a
Hermes-attributed version of this story. 64 new entries remain.
Adds 65 user stories gathered from Reddit and X, prepended to the front of
the list so the newest entries surface first.
Every quote was re-fetched from its live source and programmatically verified
to be a verbatim substring of the source text. No existing entries are
modified, reordered, or removed - this is a pure prepend.
45 Reddit, 20 X. All categories and sources already exist in the file; no
schema change.
Per the 'when in doubt, optional' rule — plan-interview is an
on-request capability, not a weekly daily-driver for most users.
Install via: hermes skills install official/software-development/grill-me