Files
hermes-agent/hermes_cli/banner.py
T
Siddharth Balyan 4bdd64b334 The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract

The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an
anonymous account exactly one model on its own host, refuses everything else with a structured
429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and
tells a signed-in account that still asks for `nous/welcome` what to switch to in an
`x-nous-model-switch` header. Four client-side gaps against that contract:

- Auxiliary calls were refused on every session. The auxiliary client asked the welcome host
  for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free`
  before each fallback. On the welcome host it now uses `nous/welcome` (its backing model
  covers auxiliary work) and skips Nous for vision, which the welcome model does not take.

- The structured 429 body was never read. The classifier now parses `reason` /
  `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are
  non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited`
  are rate limits that honour `retry_after` and never rotate the free tier's only credential.
  The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and
  fall back instead of retrying or re-exchanging. The terminal paths say what happened and
  name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal).

- The `x-nous-model-switch` header was ignored. The chat-completions transport records it
  beside the rate-limit and credits headers; the next call moves the session, and the config
  default when it still names `nous/welcome`, to the backing model the gateway named.

- A guest fell back to the paid host. With `inference_base_url` absent from the exchange or
  outside the host allowlist, routing defaulted to inference-api, where every request is a
  400. A guest now defaults to the welcome literal at the exchange, in the shared store's
  shape, and in effective routing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8)

* feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use

A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided
setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier
picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime
rung now sets it up there instead of failing "not logged in", so the guided chat no longer races
the root profile's first-run mint.

nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever
nothing else is configured; "on-request" mints only when the free tier is asked for by name
(nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers —
the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the
connector token path — still adopt what the shared store holds, so every profile follows the one
identity the guided setup created, but never create one on their own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c)
(cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165)

* feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit

Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity
(nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is
configured. "explicit" means Hermes never creates one on its own: the only creator is the new
provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup
on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and
before the guided chat exists — so the identity lands in the root store every profile reads
through and is there before any session asks for nous/welcome. That closes the race against the
backend's own setup, and makes "only when the setup-bot flow is used" literally true.

The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome
(the hermes model row, --provider nous), which treated a model name as intent and was broader
than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not
logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to
sign in from. Implicit callers still adopt an identity the shared store holds, and a retired
credential is replaced (a continuation, not a creation).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0)

* fix(auth): remove the nous.guest_setup knob; the free tier is created on first use

`nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its
default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`,
unknown values read as `auto`, and under `explicit` a CLI-only install could never get an
identity, which contradicts the first-run contract (first command mints, then chats).

The mint race the knob accompanied is already benign: every caller takes the profile lock then
the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup
win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which
stays. `nous.guest` remains the only free-tier policy.

Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through
`ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call
sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config
entries. The three policy tests that hold regardless of the knob are kept under
`TestExplicitProvision`; the two that only tested the knob are deleted.

(cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261)

* fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent

When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway
serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch`
moves the live session to that model and moves `config.yaml`'s default off the alias in the
same step. The messaging gateway's fallback-eviction check compares the agent's model with the
config default and evicts on any mismatch that is not a /model override, so when the config
write did not land (unreadable config, lock) the cached agent was evicted once per turn, and
prompt caching with it.

`apply_model_switch` now stamps the alias it moved the session off on the agent, and
`_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as
deliberate, beside the existing /model override case. The check takes the agent and the config
model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them.

(cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954)

* fix(auth): the free tier outranks implicit host credentials in provider resolution

On a fresh install with a leftover ~/.aws profile, resolve_provider("auto")
reached the Bedrock rung before the free-tier rung, so the first turn ran on
Bedrock and failed 403 while the free tier was still being minted in the
background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s,
three retries, no answer; the next process then switched to nous/welcome.

The free-tier rung now sits directly above the Bedrock chain: when nous.guest
is on, an existing free-tier identity answers, else a blocking mint runs, and
only then does the boto chain get a say. Everything above is unchanged and
still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter
pool, a logged-in active_provider. nous.guest: false skips the rung, and a
failed mint still falls through to Bedrock and the no-provider guidance.

Tests: six precedence cases (identity present, fresh mint, free tier off, env
key still wins, sign-in still wins, failed mint falls through). The opt-out
test now neutralizes the AWS chain like the precedence tests do; on a machine
with ~/.aws it was failing for the same reason as the bug.

Live after the fix, same Mac, AWS credentials visible, isolated shared store:
identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in
11 s.

(cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a)

* fix(auth): review follow-ups for the free-tier rung (NS-829)

- tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches
  the free tier off; its contract is the boto chain, and the free tier now
  sits above it.
- gateway/run_notifications.py: the free-tier startup line reads auth.json
  before consulting the resolver, so a gateway boot on a machine with AWS
  credentials never mints or refreshes over the network.
- hermes_cli/anon_auth.py: module docstring says where the free tier sits in
  the ladder instead of "the ladder is untouched".
- tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of
  six (parametrized ladder cases; a failed mint that returns None or raises
  falls through to Bedrock).

scripts/run_tests.sh on the five affected files: 147 passed, 0 failed.

(cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11)

* feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone

The free tier is pre-GA. Until GA it must not exist for anyone who did not
ask for it: no identity minted, no portal traffic, no free-tier copy on any
surface. One environment variable now decides that, and one function reads it.

`guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly
"1"; only then does `nous.guest` (the user's off switch) get consulted. Every
free-tier site already funnels through `guest_enabled()`, so the gate closes
minting, routing, connector entitlement, status lines and the picker row in
one place. With the variable unset, `resolve_provider("auto")` on a fresh
install raises `no_provider_configured` exactly as upstream does.

`HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the
gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted
identities as a side effect of provider resolution, and `_has_any_provider_
configured` read them ahead of every other check, making the CLI a second
reader of a flag that must have exactly one. `_forced_new_done` and the
`force` parameter of `_reconcile_and_provision` go with them.

Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in
b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847).

Not a user preference: the variable is never written to config.yaml or .env
and never shown in setup. It is deleted at GA together with its comment in
anon_auth.py. This is a deliberate, temporary exception to the "no new
HERMES_* env vars for non-secret config" rule.

Tests: fixtures set the gate instead of deleting the old lever; one new
invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "",
"0", "true" and "new" all leave the tier off with zero portal calls, red on the
previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the
feature.

* feat(auth): the free-tier identity is created in one place, at boot; every other site is a read

Before this commit eight sites could create a Nous free-tier identity as a
side effect of something else: resolving a provider, the CLI's first-run
check, the CLI's session setup (in the background beside an own key), a
connector bearer read, the desktop polling `free_tier.status`, the sign-in
precondition, the desktop's `free_tier.provision`, and the dead-credential
re-mint. A poll could mint. Provider resolution could hit the network. Two
of them raced each other on a fresh install.

Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator.
`hermes serve` runs it on a daemon thread from `_lifespan` beside the other
background boots; `cmd_chat` runs it synchronously before the first-run
guard. It inventories credentials first (`resolve_provider("auto",
skip_free_tier=True)`: what would carry inference if the free tier did not
exist), creates the identity only when `guest_enabled()`, resolves inference,
records a `SetupRecord` in process memory and broadcasts ONE `setup.ready`
event. It runs on every boot; only the mint is gated.

`ensure_portal_identity` now requires `explicit=True` and raises otherwise.
Its callers are the bootstrap, the desktop's `free_tier.provision` (the
explicit retry when the boot could not create the identity) and the two
dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`,
`managed_tool_gateway._replace_dead_guest_token`). The background thread
path and `provision_free_tier` are deleted with their last callers.

Reads that used to mint and now only read: `auth.py::resolve_provider`
rung 7 (an existing identity still outranks the Bedrock chain, NS-829
ordering kept), `main.py::_has_any_provider_configured`,
`cli_agent_setup_mixin._ensure_runtime_credentials`,
`managed_tool_gateway.read_nous_access_token` (no identity -> None),
`anon_sign_in.run_sign_in` (no identity -> Unavailable),
`methods_free_tier` `free_tier.status`.

`setup.status` answers from the record for the launch profile, blocking up
to 8 s while the bootstrap is in flight so a client's first poll lands after
the identity exists rather than racing it; a named profile, or a process
that never ran the bootstrap, keeps today's live probe. The record's fields
ride along additively (`ready`, `free_tier`, `other_providers`,
`inference_provider`).

Identity and inference are decoupled (NS-845 Q1.3): the mint sets
`active_provider="nous"` only when the inventory found nothing else usable
(`_mint_locked(carries_inference=)`); an adopted account always does. A token
refresh no longer re-elects the provider it refreshed
(`_save_provider_state_to_source` writes credentials, not the user's
choice) — that write was how an own-key install ended up on the free tier
after the first connector call.

Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check),
bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint),
62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and
`provision_free_tier`), and a04b05260c (blocking mint in the resolver).
Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847.

Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps
inference; reads never reach the portal; a refused mint is memoised),
`free_tier.status` fails loudly if it ever calls the creator, the resolver
stub fails loudly if resolution ever mints, `setup.status` reads the record,
`skip_free_tier` proves the inventory question. The three sign-in tests for
the deleted pre-mint collapse into one (`no identity -> Unavailable, zero
portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and
off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI
matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20.

* fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup"

A free-tier identity carries $0 by design, so the portal seed reports
`paid_access=False` for it. `is_free_tier_model` did not know the welcome
host, read that as a depleted account, and every free-tier turn ended with
the credits-depleted notice telling the user to top up an account they do
not have.

Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host
(`anon_auth.route_is_welcome_host`) is the free tier. The host is the
evidence, not the model name: the paid inference host can serve
`nous/welcome` to a named account and that account's depletion is real, so
`("nous/welcome", <inference host>)` stays False. Local data only, like the
three rules above it.

Restores the two contracts dropped by hermes-magic 674e11d1eaa (the
prototype line ran without unit tests): the welcome host is free without
any pricing evidence; the model name alone is not. The first is red without
this fix.

* fix(copy): free-tier text stops promising a connector transfer and never names the config key

Sign-in copy on every surface said "Sign in to keep your connectors" and
ended with "Your connectors are kept." The transfer registry that would
make that true is empty (NS-821): nothing carries over today. The copy now
says what signing in does give ("unlock more models and tools") and the
completion line names the account, not a transfer. The docs page loses the
"connectors carry over" paragraph for the same reason.

The picker's off-state line exposed `nous.guest: false` and the word
"guest"; user copy names the free tier only (R-USR-1).

The docs page gains the pre-rollout note: until GA nothing on it happens
without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and
"replaced on next use" sentences now describe the boot bootstrap.

zh is a strict locale: the `freeTier` block was English placeholder text
copied from `en`; it is now Chinese. `connectorsKept` is renamed
`completedBody` since it no longer talks about connectors.

* feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn

The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1"
as on. Until now nothing in the desktop set it, so a packaged app could
never turn the free tier on, and a backend spawned by the app could
disagree with the app about whether the tier was live.

`electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled`
is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has
`--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch
into a module constant. `desktopBackendSpawnEnv` wraps every backend env
as the outermost call and writes the flag LAST, as "1" or an explicit "0",
so no earlier spread (`process.env`, `backend.env`) can resurrect a stray
value from the parent shell.

Stamped onto all three spawn sites: the primary `serve` spawn, the pooled
per-profile spawn, and the remote SSH `exec env ...` command (which gains
` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and
the backend probes are not backend spawns and do not get it: a
`hermes --tui` typed in the pane must not mint.

The renderer learns the same fact read-only through the existing
`hermes:launch-flags` sync IPC (`guestOnboarding`) and preload
(`window.hermesDesktop.guestOnboardingEnabled`).

Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding`
maps to it in main). Two invariant tests on the pure helpers: only "1" or
the argv flag enables; the spawn env carries "1"/"0" as the last word and
preserves every other key.

* feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll

The backend's boot bootstrap now announces `setup.ready` once, after it has
created (or refused) the free-tier identity and resolved the inference
route. The renderer used to discover both by polling `setup.status`,
`setup.runtime_check` and `free_tier.status` every 60 s from
`useStatusSnapshot`; a fresh install's chip, notice strip and onboarding
overlay could sit stale for up to a minute after boot, and three RPCs a
minute per window kept asking a question whose answer changes only at
boundaries the backend already announces.

`handleLifecycleEvent` routes `setup.ready` (active source only, like
`skin.changed`) to `notifySetupReady()`, a one-shot tick atom in
`live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens
to it and runs one readiness round at once (`setup.status` +
`setup.runtime_check` + `free_tier.status`). The readiness legs also run
once on open and on return from another app, as today. The 60 s tick keeps
only `getStatus()`.

`SetupStatusSnapshot` types the record's additive fields (`ready`,
`free_tier`, `other_providers`, `inference_provider`); readiness semantics
are unchanged and still key on `provider_configured` + `runtime_check`.

Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one
refresh from the active source and none from another; the snapshot hook's
contract is three legs on open, one leg on the tick.

* fix(cli): the banner names the free tier's model instead of "no model configured"

The welcome banner prints before credentials resolve, so on a fresh install
`model` is empty and the banner said, in red, "no model configured — run
/model or hermes setup". Under the free tier that is false: the route is
already known from local state (identity on disk, tier on), and the first
message will run on `nous/welcome`.

`_banner_left_lines` now asks the route the same question when `model` is
empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous
Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the
banner's 'no model configured' line reads the resolved route").

Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`;
gate off -> the red line, zero portal calls.

* fix(aux): vision on the free tier uses nous/welcome too

The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the
backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the
repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the
welcome host would have sent every image step past the free tier for no reason, so the
auxiliary client pins the route's one model for every lane. A backing model that takes
no images answers with the upstream's own error, which the ladder handles as it always has.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415)

* fix(gateway): hermes gateway run is a boot owner of the free tier too

Rung 5 made every demand-time free-tier site a read: resolve_provider,
the connector token, the /login precondition. That is only correct if
every process that can reach those sites ran the bootstrap first. The
CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone
messaging gateway did not. A fresh HERMES_HOME with the gate on and
`hermes gateway run` reached provider resolution with no identity to
consume, and /login returned Unavailable. Reported by @andrexibiza on
#107697 (P1).

GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an
executor thread right after startup recovery and BEFORE any adapter
connects, so a fast first DM cannot arrive with nothing to resolve. It
is its own step, not part of the turn-machinery warm-up: the warm-up is
an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0);
the bootstrap is correctness and must always run. With the gate unset it
is a local inventory and no network.

Live, real GatewayRunner.start against a fake portal in a fresh home:
  gate on   -> 1 create, identity persisted, resolve_runtime_provider=nous,
               /login precondition sees the identity
  gate off  -> 0 portal calls, no identity, no_provider_configured
Before the fix the gate-on row was identical to the gate-off row.

Test: the bootstrap seam runs before _start_prefilter_platforms and
delegates to the one creator. Red on 5554eb6993 (no seam), green here.

---------

Co-authored-by: Robin Fernandes <robin@soal.org>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 03:45:33 +05:30

933 lines
45 KiB
Python
Raw Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Welcome banner, ASCII art, skills summary, and update check for the CLI."""
import json
import logging
import os
import shutil
import subprocess
import threading
import time
from pathlib import Path
from urllib.parse import urlparse
from hermes_constants import get_hermes_home
from typing import TYPE_CHECKING, Any, Dict, List, Optional
# rich and prompt_toolkit are imported lazily: this module sits on the TUI gateway's critical
# startup path purely for the lightweight update-check helpers, and eager rich/prompt_toolkit
# imports cost ~50ms before ``gateway.ready`` could fire.
if TYPE_CHECKING:
from rich.console import Console
logger = logging.getLogger(__name__)
# ANSI building blocks for conversation display (``_DIM``/``_RST`` are imported by callbacks.py).
_DIM = "\033[2m"
_RST = "\033[0m"
def _quiet(fn, default=None):
"""``fn()``, or ``default`` on any exception — for best-effort display inputs."""
try:
return fn()
except Exception:
return default
def cprint(text: str):
"""Print ANSI-colored text through prompt_toolkit's renderer."""
from prompt_toolkit import print_formatted_text as _pt_print
from prompt_toolkit.formatted_text import ANSI as _PT_ANSI
# prompt_toolkit needs a real console: on Windows a redirected/absent stdout raises
# NoConsoleScreenBufferError, and display helpers must never crash the caller over that.
if _quiet(lambda: _pt_print(_PT_ANSI(text)) or True) is None:
print(text)
def _active_skin():
"""The active skin object (raises when the skin engine is unavailable)."""
from hermes_cli.skin_engine import get_active_skin
return get_active_skin()
def _skin_color(key: str, fallback: str) -> str:
"""Get a color from the active skin, or return fallback."""
return _quiet(lambda: _active_skin().get_color(key, fallback), fallback)
# === ASCII Art & Branding ===
from hermes_cli import __version__ as VERSION, __release_date__ as RELEASE_DATE
HERMES_AGENT_LOGO = """[bold #FFD700]██╗ ██╗███████╗██████╗ ███╗ ███╗███████╗███████╗ █████╗ ██████╗ ███████╗███╗ ██╗████████╗[/]
[bold #FFD700]██║ ██║██╔════╝██╔══██╗████╗ ████║██╔════╝██╔════╝ ██╔══██╗██╔════╝ ██╔════╝████╗ ██║╚══██╔══╝[/]
[#FFBF00]███████║█████╗ ██████╔╝██╔████╔██║█████╗ ███████╗█████╗███████║██║ ███╗█████╗ ██╔██╗ ██║ ██║[/]
[#FFBF00]██╔══██║██╔══╝ ██╔══██╗██║╚██╔╝██║██╔══╝ ╚════██║╚════╝██╔══██║██║ ██║██╔══╝ ██║╚██╗██║ ██║[/]
[#CD7F32]██║ ██║███████╗██║ ██║██║ ╚═╝ ██║███████╗███████║ ██║ ██║╚██████╔╝███████╗██║ ╚████║ ██║[/]
[#CD7F32]╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝╚═╝ ╚═╝╚══════╝╚══════╝ ╚═╝ ╚═╝ ╚═════╝ ╚══════╝╚═╝ ╚═══╝ ╚═╝[/]"""
HERMES_CADUCEUS = """[#CD7F32]⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⢀⣀⡀⠀⣀⣀⠀⢀⣀⡀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀[/]
[#CD7F32]⠀⠀⠀⠀⠀⠀⢀⣠⣴⣾⣿⣿⣇⠸⣿⣿⠇⣸⣿⣿⣷⣦⣄⡀⠀⠀⠀⠀⠀⠀[/]
[#FFBF00]⠀⢀⣠⣴⣶⠿⠋⣩⡿⣿⡿⠻⣿⡇⢠⡄⢸⣿⠟⢿⣿⢿⣍⠙⠿⣶⣦⣄⡀⠀[/]
[#FFBF00]⠀⠀⠉⠉⠁⠶⠟⠋⠀⠉⠀⢀⣈⣁⡈⢁⣈⣁⡀⠀⠉⠀⠙⠻⠶⠈⠉⠉⠀⠀[/]
[#FFD700]⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⣴⣿⡿⠛⢁⡈⠛⢿⣿⣦⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀[/]
[#FFD700]⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠿⣿⣦⣤⣈⠁⢠⣴⣿⠿⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀[/]
[#FFBF00]⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠈⠉⠻⢿⣿⣦⡉⠁⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀[/]
[#FFBF00]⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠘⢷⣦⣈⠛⠃⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀[/]
[#CD7F32]⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⢠⣴⠦⠈⠙⠿⣦⡄⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀[/]
[#CD7F32]⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠸⣿⣤⡈⠁⢤⣿⠇⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀[/]
[#B8860B]⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠉⠛⠷⠄⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀[/]
[#B8860B]⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⢀⣀⠑⢶⣄⡀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀[/]
[#B8860B]⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⣿⠁⢰⡆⠈⡿⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀[/]
[#B8860B]⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠈⠳⠈⣡⠞⠁⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀[/]
[#B8860B]⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠈⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀[/]"""
# === Skills scanning ===
# Per-process caches: ``None`` until computed, then a 1-tuple ``(value,)`` so a computed ``None``
# is distinguishable from "not yet computed". Reset by assigning ``None`` (tests, ``hermes skills``).
_available_skills_cache: Optional[tuple] = None
_git_banner_state_cache: Optional[tuple] = None
_latest_release_cache: Optional[tuple] = None
_UNCACHED = object() # compute() result that must not be memoized
def _memo(cache_name: str, compute):
"""Return the cached value under module global ``cache_name``, computing (and storing) it once."""
cached = globals()[cache_name]
if cached is not None:
return cached[0]
value = compute()
if value is not _UNCACHED:
globals()[cache_name] = (value,)
return value
def get_available_skills() -> Dict[str, List[str]]:
"""Return skills grouped by category, filtered by platform and disabled state.
Cached per-process (the skills-tree walk costs ~100ms and feeds only the startup banner);
``prefetch_banner_data()`` pays it off-thread. A failed scan yields ``{}`` and is not cached.
"""
def _scan():
from tools.skills_tool import _find_all_skills
return _find_all_skills() # already filtered
def _compute():
all_skills = _quiet(_scan)
if all_skills is None:
return _UNCACHED
skills_by_category: Dict[str, List[str]] = {}
for skill in all_skills:
skills_by_category.setdefault(skill.get("category") or "general", []).append(skill["name"])
return skills_by_category
result = _memo("_available_skills_cache", _compute)
return {} if result is _UNCACHED else result
# === Update check ===
_UPDATE_CHECK_CACHE_SECONDS = 6 * 3600 # avoid repeated git fetches
# Returned when an update is known to exist but commits can't be counted (e.g. nix builds).
UPDATE_AVAILABLE_NO_COUNT = -1
_UPSTREAM_REPO_URL = "https://github.com/NousResearch/hermes-agent.git"
_OFFICIAL_REPO_CANONICAL = "github.com/nousresearch/hermes-agent"
def _canonical_github_remote(url: str | None) -> str:
"""Return ``host/owner/repo`` for common GitHub remote URL forms."""
if not url:
return ""
value = url.strip()
for ssh_prefix in ("git@github.com:", "ssh://git@github.com/"):
if value.startswith(ssh_prefix):
value = "github.com/" + value[len(ssh_prefix):]
break
else:
parsed = urlparse(value)
if parsed.netloc and parsed.path:
value = f"{parsed.netloc}{parsed.path}"
return value.strip().rstrip("/").removesuffix(".git").lower()
def _is_official_ssh_remote(url: str | None) -> bool:
return bool(url) and url.strip().lower().startswith(("git@", "ssh://")) and (
_canonical_github_remote(url) == _OFFICIAL_REPO_CANONICAL)
_GIT_TEXT_KW = {"text": True, "encoding": "utf-8", "errors": "replace"}
def _git_run(args: list[str], *, cwd: Optional[Path] = None, timeout: int = 5, text: bool = True,
network: bool = False):
"""Run ``git <args>`` with the shared subprocess boilerplate; None on any exception.
git output is UTF-8; on Windows ``text=True`` defaults to the ANSI code page and a byte like the
3rd of 🐛 in a commit subject crashes the stdlib reader thread (#52649), hence the explicit
encoding. ``network=True`` (ls-remote/fetch) detaches stdin and disables git/GCM prompts so a
passive update check can never hang on a ``Username for 'https://github.com':`` prompt.
"""
from hermes_cli._subprocess_compat import noninteractive_git_env, windows_hide_flags
# The banner/update probes run from GUI-hosted backends too (desktop-spawned
# ``hermes serve``), where a bare git child flashes a console window.
kwargs: dict = {"creationflags": windows_hide_flags()}
if network:
kwargs.update({"stdin": subprocess.DEVNULL, "env": noninteractive_git_env()})
try:
return subprocess.run(
["git", *args], capture_output=True, timeout=timeout, cwd=str(cwd) if cwd is not None else None,
**(_GIT_TEXT_KW if text else {}), **kwargs)
except Exception:
return None
def _git_stdout(args: list[str], *, cwd: Path, timeout: int = 5, network: bool = False) -> Optional[str]:
result = _git_run(args, cwd=cwd, timeout=timeout, network=network)
if result is None or result.returncode != 0:
return None
return (result.stdout or "").strip()
def _git_ok(args: list[str], **kw) -> bool:
"""True when ``git <args>`` ran and exited 0 (output discarded)."""
result = _git_run(args, text=False, **kw)
return result is not None and result.returncode == 0
def _git_count(args: list[str], *, cwd: Path) -> Optional[int]:
"""``int`` of a successful ``git rev-list --count``-style command, else None.
Deliberately bypasses ``_git_stdout`` so tests can stub the two layers independently.
"""
result = _git_run(args, cwd=cwd)
if result is not None and result.returncode == 0:
return _quiet(lambda: int(result.stdout.strip()))
return None
def _is_full_sha(value: Optional[str]) -> bool:
return isinstance(value, str) and len(value) == 40 and all(c in "0123456789abcdefABCDEF" for c in value)
def _github_compare_behind(current_rev: str, target_rev: str) -> Optional[int]:
"""Exact behind-count via the GitHub compare API for uncountable graphs.
Shallow installer clones and ls-remote-only probes know the two tip SHAs but have no local
history to run ``rev-list --count`` across.
"""
if not (_is_full_sha(current_rev) and _is_full_sha(target_rev)):
return None
url = f"https://api.github.com/repos/nousresearch/hermes-agent/compare/{current_rev}...{target_rev}"
def _fetch():
import urllib.request
# api.github.com 403s requests without a User-Agent.
req = urllib.request.Request(
url, headers={"Accept": "application/vnd.github+json", "User-Agent": "hermes-cli-update-check"})
with urllib.request.urlopen(req, timeout=10) as resp:
return json.loads(resp.read().decode("utf-8"))
payload = _quiet(_fetch)
ahead = payload.get("ahead_by") if isinstance(payload, dict) else None
return ahead if isinstance(ahead, int) and not isinstance(ahead, bool) and ahead >= 0 else None
def _tips_behind(head_rev: Optional[str], target_rev: Optional[str], repo_dir: Optional[Path] = None) -> Optional[int]:
"""Behind-count from two tip SHAs: None if either is unknown, 0 when equal, else count/sentinel.
With ``repo_dir``, a target that is already an ancestor of HEAD (local-ahead checkout) is 0 too.
``ahead_by == 0`` with differing tips means the remote tip is reachable from our HEAD — NOT
behind. A local-only HEAD 404s on the API, which degrades to ``UPDATE_AVAILABLE_NO_COUNT`` —
never a fabricated 1.
"""
if not head_rev or not target_rev:
return None
if head_rev == target_rev or (repo_dir is not None and _git_ok(
["merge-base", "--is-ancestor", target_rev, "HEAD"], cwd=repo_dir)):
return 0
counted = _github_compare_behind(head_rev, target_rev)
return counted if counted is not None else UPDATE_AVAILABLE_NO_COUNT
def _upstream_main_sha() -> Optional[str]:
"""Tip SHA of upstream main via HTTPS ls-remote (no auth, no prompts)."""
result = _git_run(["ls-remote", _UPSTREAM_REPO_URL, "refs/heads/main"], timeout=10, network=True)
if result is None or result.returncode != 0 or not result.stdout:
return None
return result.stdout.split()[0] or None
def _check_via_rev(local_rev: str) -> Optional[int]:
"""Compare an embedded git revision to upstream main via ls-remote (see ``_tips_behind``)."""
return _tips_behind(local_rev, _upstream_main_sha())
def _check_via_local_git(repo_dir: Path) -> Optional[int]:
"""Count commits behind origin/main in a local checkout."""
# Probe the origin URL under the same config-isolated env as the fetch below. A plain
# get-url applies a global url.<https>.insteadOf rewrite, so an SSH origin masquerades as
# HTTPS, the SSH-avoiding fast path is skipped — and the fetch, whose env drops global
# config (GIT_CONFIG_GLOBAL=/dev/null), dials the raw SSH origin; its host-key prompt opens
# /dev/tty directly and steals the CLI's keystrokes (#104591).
origin_url = _git_stdout(["remote", "get-url", "origin"], cwd=repo_dir, network=True)
if _is_official_ssh_remote(origin_url):
head_rev = _git_stdout(["rev-parse", "HEAD"], cwd=repo_dir)
if not head_rev:
return None
# Passive probe via HTTPS ls-remote (never SSH — no hardware-key prompts). Tip SHAs alone
# can't distinguish "behind" from a local commit AHEAD of origin/main, and misreporting an
# ahead checkout nudges the user into `hermes update`, which can wipe carried work — hence
# the ancestor check, against the FRESH upstream SHA (a stale tracking ref can't fake an
# up-to-date report).
return _tips_behind(head_rev, _upstream_main_sha(), repo_dir)
# Installer checkouts are shallow (`git clone --depth 1`): a plain `git fetch` would unshallow
# the repo and `rev-list --count HEAD..origin/main` would report a bogus "12492 commits
# behind". Fetch with --depth 1 to preserve the boundary and compare tip SHAs instead. Full
# clones keep the exact count path. Mirrors apps/desktop/electron/main.cjs.
is_shallow = _git_stdout(["rev-parse", "--is-shallow-repository"], cwd=repo_dir) == "true"
def _fetch() -> bool:
# Self-heal abandoned git lock files first. A stale .git/shallow.lock from a crashed fetch
# makes every fetch fail silently and stale refs get compared against HEAD until a human
# removes the lock. This passive check is also the main tmp_pack GENERATOR on flaky lines,
# so it must be the janitor too (#93732).
from hermes_cli.gitlock import clear_stale_git_locks, clear_stale_tmp_packs
clear_stale_git_locks(repo_dir)
clear_stale_tmp_packs(repo_dir)
# Scope the fetch to the one branch compared against: an unscoped ``git fetch origin``
# transfers ~1,400 remote heads (3.0 s vs 0.55 s measured) and can burn the full timeout.
# A scoped fetch still updates ``origin/main`` and FETCH_HEAD; ``--depth 1`` preserves
# the shallow boundary.
fetch_args = ["fetch", "origin", "main", *(["--depth", "1"] if is_shallow else []), "--quiet"]
return _git_ok(fetch_args, cwd=repo_dir, timeout=10, network=True)
fetch_ok = _quiet(_fetch, False) # Offline or timeout — don't use stale refs
# When the fetch fails the local origin/main ref is stale: it cannot prove *currentness*, but
# if it already shows HEAD behind, that is sound evidence an update exists. Return the positive
# stale count; None (inconclusive) otherwise so the caller doesn't cache a false "up to date".
if is_shallow:
# (#82166, review #92578)
if not fetch_ok:
return None
# No history across the shallow boundary. `origin/main` may not be a tracking ref in a
# `clone --depth 1`, so prefer FETCH_HEAD (just updated) and fall back to origin/main.
head_rev = _git_stdout(["rev-parse", "HEAD"], cwd=repo_dir)
target_rev = (
_git_stdout(["rev-parse", "FETCH_HEAD"], cwd=repo_dir)
or _git_stdout(["rev-parse", "origin/main"], cwd=repo_dir))
return _tips_behind(head_rev, target_rev)
behind = _git_count(["rev-list", "--count", "HEAD..origin/main"], cwd=repo_dir)
return behind if fetch_ok or (behind is not None and behind > 0) else None
def _read_json(path: Path) -> Optional[dict]:
"""Parse ``path`` as a JSON object; None when missing, unreadable, or not a dict."""
blob = _quiet(lambda: json.loads(path.read_text(encoding="utf-8")))
return blob if isinstance(blob, dict) else None
def check_for_updates(*, passive: bool = False) -> Optional[int]:
"""Check whether a Hermes update is available.
If ``HERMES_REVISION`` is set (nix builds embed it), compare it to upstream main via
``git ls-remote``; otherwise count commits behind ``origin/main`` in the local checkout.
"""
def _read_config_opt_out():
from hermes_cli.config import load_config
return load_config().get("updates", {}).get("check", True) is False
if passive and _quiet(_read_config_opt_out) is True:
return None
cache_file = get_hermes_home() / ".update_check"
embedded_rev = os.environ.get("HERMES_REVISION") or None
# Docker images have no working tree (the image excludes `.git`) and set no HERMES_REVISION.
# None makes both the Rich banner and the Ink badge show nothing, mirroring the dashboard's
# `/api/hermes/update/check` short-circuit so the surfaces agree.
def _install_method():
from hermes_cli.config import detect_install_method, get_project_root
return detect_install_method(get_project_root())
if _quiet(_install_method) in {"docker", "apt"}:
return None
# Cache is invalidated when the embedded rev OR installed version changed since the last check.
now = time.time()
cached = _read_json(cache_file)
if (cached is not None and now - cached.get("ts", 0) < _UPDATE_CHECK_CACHE_SECONDS
and cached.get("rev") == embedded_rev and cached.get("ver") == VERSION):
return cached.get("behind")
if embedded_rev:
behind = _check_via_rev(embedded_rev)
else:
# No checkout and no embedded revision — status can't be determined.
repo_dir = _resolve_repo_dir()
behind = _check_via_local_git(repo_dir) if repo_dir is not None else None
# Don't cache inconclusive results: None means the check could not run (typically a failed
# fetch), and caching it would suppress retries for the full 6-hour window (#82166).
if behind is not None:
_quiet(lambda: cache_file.write_text(
json.dumps({"ts": now, "behind": behind, "rev": embedded_rev, "ver": VERSION}), encoding="utf-8"))
return behind
def _resolve_repo_dir() -> Optional[Path]:
"""The active Hermes git checkout, or None if this isn't a git install.
Prefers the running code's location: ``$HERMES_HOME/hermes-agent/`` may be a stale copy
carried over by ``--clone-all``.
"""
repo_dir = Path(__file__).parent.parent.resolve()
if not (repo_dir / ".git").exists():
repo_dir = get_hermes_home() / "hermes-agent"
return repo_dir if (repo_dir / ".git").exists() else None
def get_git_banner_state(repo_dir: Optional[Path] = None) -> Optional[dict]:
"""Return upstream/local git hashes for the startup banner.
Cached per-process (default ``repo_dir`` only): 2-3 git subprocesses (~100ms) whose result
cannot change under a running CLI. The cache lets ``prefetch_banner_data()`` pay it off-thread.
"""
if repo_dir is not None:
return _compute_git_banner_state(repo_dir)
return _memo("_git_banner_state_cache", _compute_git_banner_state)
def _baked_banner_state() -> Optional[dict]:
"""Banner state from the baked build SHA (Docker image path), or None."""
def _baked():
from hermes_cli.build_info import get_build_sha
return get_build_sha(short=8)
baked = _quiet(_baked)
return {"upstream": baked, "local": baked, "ahead": 0} if baked else None
def _compute_git_banner_state(repo_dir: Optional[Path] = None) -> Optional[dict]:
repo_dir = repo_dir or _resolve_repo_dir()
if repo_dir is None:
return _baked_banner_state()
upstream, local = (_git_stdout(["rev-parse", "--short=8", rev], cwd=repo_dir) for rev in ("origin/main", "HEAD"))
if not upstream or not local:
# Live-git lookup failed (e.g. shallow clone without origin/main).
return _baked_banner_state()
ahead = _git_count(["rev-list", "--count", "origin/main..HEAD"], cwd=repo_dir) or 0
return {"upstream": upstream, "local": local, "ahead": max(ahead, 0)}
_RELEASE_URL_BASE = "https://github.com/NousResearch/hermes-agent/releases/tag"
def get_latest_release_tag(repo_dir: Optional[Path] = None) -> Optional[tuple]:
"""Return ``(tag, release_url)`` for the latest local git tag, or None (a miss is cached too).
Release URL always points at the canonical NousResearch/hermes-agent repo (forks get no link).
"""
def _compute():
rd = repo_dir or _resolve_repo_dir()
tag = _git_stdout(["describe", "--tags", "--abbrev=0"], cwd=rd, timeout=3) if rd else None
return (tag, f"{_RELEASE_URL_BASE}/{tag}") if tag else None
return _memo("_latest_release_cache", _compute)
def format_banner_version_label() -> str:
"""Return the version label shown in the startup banner title."""
base = f"Hermes Agent v{VERSION} ({RELEASE_DATE})"
state = get_git_banner_state()
if not state:
return base
upstream, local = state["upstream"], state["local"]
ahead = int(state.get("ahead") or 0)
if ahead <= 0 or upstream == local:
return f"{base} · upstream {upstream}"
return f"{base} · upstream {upstream} · local {local} (+{ahead} carried {_plural(ahead, 'commit')})"
# === Non-blocking update check ===
_update_result: Optional[int] = None
_update_check_done = threading.Event()
def _daemon(name: Optional[str], target) -> None:
"""Start a daemon thread running ``target`` with any exception swallowed."""
threading.Thread(target=lambda: _quiet(target), name=name, daemon=True).start()
def prefetch_update_check():
"""Kick off update check in a background daemon thread."""
def _run():
global _update_result
_update_result = check_for_updates(passive=True)
_update_check_done.set()
_daemon(None, _run)
_banner_data_prefetch_started = False
def prefetch_banner_data():
"""Warm the banner's subprocess/I/O-heavy inputs in a daemon thread.
Git state (~130ms) and the skills index (~110ms) are cached per-process by their own modules,
so warming them while the main thread pays the CPU-bound imports overlaps GIL-releasing I/O
with import work. Idempotent; failures don't matter because the banner recomputes anything missing.
"""
global _banner_data_prefetch_started
if _banner_data_prefetch_started:
return
_banner_data_prefetch_started = True
_daemon("banner-data-prefetch", lambda: [_quiet(warm) for warm in (
get_git_banner_state, get_latest_release_tag, get_available_skills)])
def get_update_result(timeout: float = 0.5) -> Optional[int]:
"""Get result of prefetched check. Returns None if not ready."""
_update_check_done.wait(timeout=timeout)
return _update_result
def _format_update_notice(behind: int) -> str:
"""Render the update warning line for a non-zero ``behind`` result."""
from hermes_cli.config import get_managed_update_command, recommended_update_command
if behind > 0:
return (
f"[bold yellow]⚠ {behind} {_plural(behind, 'commit')} behind[/]"
f"[dim yellow] — run [bold]{recommended_update_command()}[/bold] to update[/]")
# UPDATE_AVAILABLE_NO_COUNT (nix): an update exists but we don't know by how much, nor how
# the user installed (nix run, profile, system flake, home-manager).
managed_cmd = get_managed_update_command()
suffix = f"[dim yellow] — run [bold]{managed_cmd}[/bold][/]" if managed_cmd else ""
return f"[bold yellow]⚠ update available[/]{suffix}"
_deferred_update_notice_started = False
def _render_markup_to_ansi(markup: str) -> str:
"""Rich markup → ANSI string, for output that must go through prompt_toolkit's renderer.
Under ``patch_stdout`` (the interactive CLI), a plain ``Console.print`` writes ESC bytes into
the StdoutProxy, which sanitizes them into visible ``?[1;33m…`` artifacts (#83969).
"""
from io import StringIO
from rich.console import Console as _Console
buf = StringIO()
_Console(file=buf, force_terminal=True, color_system="truecolor", highlight=False).print(markup)
return buf.getvalue().rstrip("\n")
def _defer_update_notice(max_wait: float = 30.0) -> None:
"""Print the update warning once the prefetched check completes (at most once per process).
Used when the banner rendered before the update prefetch finished so startup never blocks on
git/network. The notice lands after prompt_toolkit owns the terminal, so it is routed through
``cprint`` (prompt_toolkit's renderer prints above a running application from any thread).
"""
global _deferred_update_notice_started
if _deferred_update_notice_started:
return
_deferred_update_notice_started = True
def _wait_and_print() -> None:
if _update_check_done.wait(timeout=max_wait) and _update_result:
cprint(_render_markup_to_ansi(_format_update_notice(_update_result)))
_daemon("update-notice", _wait_and_print) # never break the session over an update notice
# === Welcome banner ===
def _plural(n: int, word: str) -> str:
return word if n == 1 else f"{word}s"
def _format_context_length(tokens: int) -> str:
"""Format a token count for display (e.g. 128000 → '128K', 1048576 → '1M')."""
for unit, div in (("M", 1_000_000), ("K", 1_000)):
if tokens >= div:
val = tokens / div
rounded = round(val)
return f"{rounded}{unit}" if abs(val - rounded) < 0.05 else f"{val:.1f}{unit}"
return str(tokens)
def _display_toolset_name(toolset_name: str) -> str:
"""Normalize internal/legacy toolset identifiers for banner display."""
return toolset_name.removesuffix("_tools") if toolset_name else "unknown"
def _short_label(name: str) -> str:
"""Truncate a model/preset slug to fit the banner's left column."""
return name[:25] + "..." if len(name) > 28 else name
# === Banner snapshot — warm-launch fast path ===
# The tool panel needs the full tool registry (~0.5-0.9s cold, the largest chunk of time-to-
# banner). The list is a pure function of (config.yaml, .env, code checkout, enabled toolsets),
# so the rendered inputs are snapshotted to disk and replayed when the fingerprint matches. The
# agent's REAL tool list is still computed fresh at first message; the snapshot only feeds the
# cosmetic panel, and a background refresh (cli.show_banner) re-verifies it right after render.
_BANNER_SNAPSHOT_VERSION = 1
def _banner_snapshot_path() -> Path:
return get_hermes_home() / "cache" / "banner_snapshot.json"
def banner_snapshot_fingerprint() -> Optional[str]:
"""Fingerprint the inputs the banner tool panel depends on."""
import hashlib
def _inputs():
from hermes_cli.config import get_config_path
return (get_config_path(), get_hermes_home() / ".env")
paths = _quiet(_inputs)
if paths is None:
return None
parts = [f"v{_BANNER_SNAPSHOT_VERSION}"]
for p in paths:
st = _quiet(p.stat)
parts.append(f"{p.name}:{st.st_mtime_ns}:{st.st_size}" if st else f"{p.name}:absent")
# Code checkout: version + git HEAD when available (post-update change).
parts.append(str(VERSION))
state = get_git_banner_state()
if state:
parts.append(str(state.get("local", "")))
return hashlib.sha256("|".join(parts).encode("utf-8")).hexdigest()
def load_banner_snapshot(enabled_toolsets: List[str] = None) -> Optional[Dict[str, Any]]:
"""Return the stored banner snapshot when its fingerprint is current."""
blob = _read_json(_banner_snapshot_path())
if blob is None:
return None
fp = banner_snapshot_fingerprint()
if (not fp or blob.get("fingerprint") != fp
or blob.get("enabled_toolsets") != sorted(enabled_toolsets or [])
or not isinstance(blob.get("tools"), list)
or not all(isinstance(blob.get(k), dict)
for k in ("toolset_map", "availability", "skills_by_category"))):
return None
return blob
def save_banner_snapshot(tools: List[dict], enabled_toolsets: List[str], availability: Dict[str, Any],
toolset_map: Dict[str, str]) -> None:
"""Persist the banner tool panel inputs for next launch (best-effort)."""
fp = banner_snapshot_fingerprint()
if not fp:
return
payload = {
"fingerprint": fp,
"enabled_toolsets": sorted(enabled_toolsets or []),
"tools": [{"function": {"name": t["function"]["name"]}}
for t in tools if isinstance(t, dict) and t.get("function", {}).get("name")],
"toolset_map": toolset_map,
"availability": {
"unavailable_toolsets": availability.get("unavailable_toolsets", []),
**{k: list(availability.get(k, [])) for k in ("lazy_tools", "disabled_tools")}},
"skills_by_category": get_available_skills(),
}
def _write():
import tempfile
path = _banner_snapshot_path()
path.parent.mkdir(parents=True, exist_ok=True)
fd, tmp = tempfile.mkstemp(dir=str(path.parent), prefix=".banner_snap.")
with os.fdopen(fd, "w", encoding="utf-8") as fh:
json.dump(payload, fh)
os.replace(tmp, path)
_quiet(_write)
def compute_toolset_availability(enabled_toolsets: List[str] = None) -> Dict[str, Any]:
"""Compute ``{"unavailable_toolsets", "lazy_tools", "disabled_tools"}`` for the banner.
Split out so the result can be snapshotted and replayed without importing ``model_tools``.
"""
from model_tools import check_tool_availability, TOOLSET_REQUIREMENTS
enabled_toolsets = enabled_toolsets or []
_, unavailable_toolsets = check_tool_availability(quiet=True)
# The availability check walks the GLOBAL registry, so it includes toolsets outside this
# agent's platform set (e.g. `discord` on a CLI session) which must never surface in
# "Available Tools". Restrict to enabled toolsets; an enabled toolset with unmet deps
# legitimately shows as disabled/lazy below.
_enabled_ts = {str(t) for t in enabled_toolsets}
if _enabled_ts:
unavailable_toolsets = [
item for item in unavailable_toolsets if str(item.get("id", item.get("name", ""))) in _enabled_ts]
# Toolsets with a check_fn are lazy-initialized (e.g. honcho): unavailable at banner time
# because the check hasn't run yet, but not misconfigured.
lazy_tools, disabled_tools = set(), set()
for item in unavailable_toolsets:
is_lazy = TOOLSET_REQUIREMENTS.get(item.get("name", ""), {}).get("check_fn")
(lazy_tools if is_lazy else disabled_tools).update(item.get("tools", []))
return {"unavailable_toolsets": unavailable_toolsets, "lazy_tools": sorted(lazy_tools),
"disabled_tools": sorted(disabled_tools)}
def _mcp_server_line(srv: dict, *, dim: str, text: str) -> str:
"""One banner line for an MCP server status entry."""
name, transport = srv["name"], srv["transport"]
if srv["connected"]:
return f"[dim {dim}]{name}[/] [{text}]({transport})[/] [dim {dim}]—[/] [{text}]{srv['tools']} tool(s)[/]"
status = "disabled" if srv.get("disabled") else srv.get("status")
suffix = {"disabled": f"[dim {dim}]— disabled[/]", "connecting": "[yellow]— connecting[/]",
"configured": f"[dim {dim}]— configured[/]"}.get(status)
if suffix is not None:
return f"[dim {dim}]{name}[/] [dim]({transport})[/] {suffix}"
return f"[red]{name}[/] [dim]({transport})[/] [red]— failed[/]"
def _truncate_tool_names(tool_names: List[str]) -> List[Optional[str]]:
"""Cut a toolset's tool list to ~42 columns; ``None`` marks the elided tail."""
if len(", ".join(tool_names)) <= 45:
return list(tool_names)
short_names: List[Optional[str]] = []
length = 0
for name in tool_names:
if length + len(name) + 2 > 42:
short_names.append(None)
break
short_names.append(name)
length += len(name) + 2
return short_names
def _pack_skill_names(skill_names: List[str], avail: int) -> str:
"""Join skill names into ``avail`` columns, ending with ``+N more`` when they don't all fit."""
parts: List[str] = []
length = 0
for i, name in enumerate(skill_names):
needed = (2 if parts else 0) + len(name)
after = len(skill_names) - (i + 1) # indicator size IF we add this skill then stop
ind_len = len(f", +{after} more") if after > 0 else 0
if parts and length + needed + ind_len > avail:
parts.append(f"+{len(skill_names) - len(parts)} more")
break
parts.append(name)
length += needed
return ", ".join(parts)
def _moa_aggregator_label(preset_name: str) -> str:
"""Short aggregator-model label for a MoA preset ("" when the preset has none)."""
from hermes_cli.config import load_config
from hermes_cli.moa_config import normalize_moa_config
preset = normalize_moa_config(load_config().get("moa") or {}).get("presets", {}).get(preset_name)
model = str(((preset or {}).get("aggregator") or {}).get("model") or "")
return model.split("/")[-1]
def _mcp_configured() -> bool:
"""Cheap probe: does config.yaml or the persisted plugin key cache name any MCP server?
The full ``get_mcp_status()`` path resolves portable plugin MCP servers, which JOINS the in-flight
background plugin discovery (~100ms on the startup path), so skip it when nothing is configured.
When either probe can't tell, take the full path.
"""
def _native():
from hermes_cli.config import load_config
return bool((load_config() or {}).get("mcp_servers"))
def _portable():
from hermes_cli.plugins import get_portable_mcp_server_names_nowait
return bool(get_portable_mcp_server_names_nowait())
return _quiet(_native, True) or _quiet(_portable, True)
def _probe_mcp_status() -> list:
from tools.mcp_tool_discovery import get_mcp_status
return get_mcp_status()
def _codex_runtime_active() -> bool:
"""True when the codex_app_server runtime is active (tool counts then live inside codex)."""
from hermes_cli.codex_runtime_switch import get_current_runtime
from hermes_cli.config import load_config
return get_current_runtime(load_config()) == "codex_app_server"
def _active_profile_name() -> Optional[str]:
from hermes_cli.profiles import get_active_profile_name
return get_active_profile_name()
def _route_model_for_banner(provider: Any) -> str:
"""The model the resolved route will actually serve when config names none: today only the Nous
free tier (welcome host -> ``nous/welcome``). Read from the boot record and local auth state;
no network. Empty when nothing resolves, so the caller keeps its "no model configured" line."""
if (provider or "auto").strip().lower() not in ("auto", "nous"):
return ""
from hermes_cli.anon_auth import GUEST_MODEL, guest_carries_inference
return GUEST_MODEL if guest_carries_inference() else ""
def _banner_left_lines(model: str, cwd: str, session_id, context_length, provider, *, accent: str, dim: str) -> list:
"""Model / cwd / session lines under the hero art."""
def _dim_sep(label: str) -> str:
return f" [dim {dim}]·[/] [dim {dim}]{label}[/]"
lines = []
ctx_str = _dim_sep(f"{_format_context_length(context_length)} context") if context_length else ""
nous_str = _dim_sep("Nous Research")
if not (model or "").strip():
# Credentials resolve lazily on the first message; the banner prints first. Ask the route
# the same question so a fresh free-tier install shows its model, not a red "unconfigured".
model = _quiet(lambda: _route_model_for_banner(provider), "") or model
if (provider or "").strip().lower() == "moa":
# MoA virtual provider: ``model`` is a preset name; show it with its aggregator.
agg_label = _quiet(lambda: _moa_aggregator_label(model), "")
agg_str = _dim_sep(f"agg {agg_label}") if agg_label else ""
lines.append(f"[{accent}]MoA: {_short_label(model)}[/]{agg_str}{ctx_str}{nous_str}")
elif not (model or "").strip() or (model or "").strip().lower() == "unknown":
# Unconfigured install: the clearest place to say what is wrong and how to fix it.
lines.append(f"[bold red]no model configured[/] [dim {dim}]— run /model or hermes setup[/]")
else:
model_short = model.split("/")[-1].removesuffix(".gguf")
lines.append(f"[{accent}]{_short_label(model_short)}[/]{ctx_str}{nous_str}")
if os.getenv("HERMES_YOLO_MODE"):
lines.append(f"[bold red]⚠ YOLO mode[/] [dim {dim}]— all approval prompts bypassed[/]")
lines.append(f"[dim {dim}]{cwd}[/]")
if session_id:
lines.append(f"[dim {_skin_color('session_border', '#8B8682')}]Session: {session_id}[/]")
return lines
def _banner_tool_lines(
tools: list, unavailable_toolsets: list, get_toolset_for_tool, *,
lazy_tools: set, disabled_tools: set, accent: str, dim: str, text: str) -> list:
""""Available Tools" section: up to 8 toolsets, each truncated to ~42 columns."""
lines = [f"[bold {accent}]Available Tools[/]"]
toolsets_dict: Dict[str, list] = {}
for tool in tools:
tool_name = tool["function"]["name"]
toolset = _display_toolset_name(get_toolset_for_tool(tool_name) or "other")
toolsets_dict.setdefault(toolset, []).append(tool_name)
for item in unavailable_toolsets:
names = toolsets_dict.setdefault(_display_toolset_name(item.get("id", item.get("name", "unknown"))), [])
for tool_name in item.get("tools", []):
if tool_name not in names:
names.append(tool_name)
def _color_tool(name: Optional[str]) -> str:
if name is None: # truncation marker
return "[dim]...[/]"
color = "red" if name in disabled_tools else "yellow" if name in lazy_tools else text
return f"[{color}]{name}[/]"
sorted_toolsets = sorted(toolsets_dict.keys())
for toolset in sorted_toolsets[:8]:
tool_names = _truncate_tool_names(sorted(toolsets_dict[toolset]))
lines.append(f"[dim {dim}]{toolset}:[/] {', '.join(_color_tool(n) for n in tool_names)}")
if len(sorted_toolsets) > 8:
lines.append(f"[dim {dim}](and {len(sorted_toolsets) - 8} more toolsets...)[/]")
return lines
def _banner_skill_lines(skills_by_category: Dict[str, List[str]], skills_enabled: bool, *, dim: str, text: str) -> list:
""""Available Skills" body, sized to ~60% of the terminal width (the right grid column)."""
if not skills_enabled:
return [f"[dim {dim}]Skills toolset disabled[/]"]
if not skills_by_category:
return [f"[dim {dim}]No skills installed[/]"]
right_col_width = max(int(shutil.get_terminal_size().columns * 0.6) - 10, 30)
lines = []
for category in sorted(skills_by_category.keys()):
# Account for the "category: " prefix.
skills_str = _pack_skill_names(sorted(skills_by_category[category]), max(right_col_width - len(category) - 2, 20))
lines.append(f"[dim {dim}]{category}:[/] [{text}]{skills_str}[/]")
return lines
def build_welcome_banner(
console: "Console", model: str, cwd: str, tools: List[dict] = None, enabled_toolsets: List[str] = None,
session_id: str = None, get_toolset_for_tool=None, context_length: int = None, provider: str = None,
availability: Dict[str, Any] = None, skills_by_category: Dict[str, List[str]] = None,
):
"""Build and print a welcome banner with caduceus on left and info on right.
When ``provider == "moa"``, ``model`` is a MoA preset name and the aggregator is rendered.
Passing a precomputed ``availability`` together with ``get_toolset_for_tool`` avoids any
``model_tools`` import (banner snapshot replay).
"""
from rich.panel import Panel
from rich.table import Table
if get_toolset_for_tool is None:
from model_tools import get_toolset_for_tool
tools = tools or []
enabled_toolsets = enabled_toolsets or []
if availability is None:
availability = compute_toolset_availability(enabled_toolsets)
_enabled_ts = {str(t) for t in enabled_toolsets}
# Resolve skin colors once for the entire banner
accent = _skin_color("banner_accent", "#FFBF00")
dim = _skin_color("banner_dim", "#B8860B")
text = _skin_color("banner_text", "#FFF8DC")
# Use skin's custom caduceus art if provided
_bskin = _quiet(_active_skin)
left_lines = ["", getattr(_bskin, "banner_hero", None) or HERMES_CADUCEUS, ""]
left_lines += _banner_left_lines(model, cwd, session_id, context_length, provider, accent=accent, dim=dim)
right_lines = _banner_tool_lines(
tools, availability.get("unavailable_toolsets", []), get_toolset_for_tool,
lazy_tools=set(availability.get("lazy_tools", [])), disabled_tools=set(availability.get("disabled_tools", [])),
accent=accent, dim=dim, text=text)
# MCP Servers section (only if configured) — see ``_mcp_configured`` for why the cheap probe.
mcp_status = _quiet(_probe_mcp_status, []) if _mcp_configured() else []
if mcp_status:
right_lines += ["", f"[bold {accent}]MCP Servers[/]"]
right_lines.extend(_mcp_server_line(srv, dim=dim, text=text) for srv in mcp_status)
right_lines += ["", f"[bold {accent}]Available Skills[/]"]
# The skills catalog is only reachable when the `skills` toolset is enabled (skill_view /
# skill_manage). When disabled (Blank Slate) the agent cannot load any skill, so advertising
# the on-disk catalog would be misleading — reflect the real state.
_skills_enabled = (not _enabled_ts) or ("skills" in _enabled_ts)
if not _skills_enabled:
skills_by_category = {}
elif skills_by_category is None:
skills_by_category = get_available_skills()
total_skills = sum(len(s) for s in skills_by_category.values())
right_lines += _banner_skill_lines(skills_by_category, _skills_enabled, dim=dim, text=text)
right_lines.append("")
mcp_connected = sum(1 for s in mcp_status if s["connected"])
summary_parts = [f"{len(tools)} tools", f"{total_skills} skills"]
if mcp_connected:
summary_parts.append(f"{mcp_connected} MCP servers")
summary_parts.append("/help for commands")
# Flag the codex_app_server runtime so users understand why tool counts may not match what's
# reachable (codex builds its own tool list inside the spawned subprocess).
if _quiet(_codex_runtime_active, False):
right_lines.append(f"[bold {accent}]Runtime:[/] [{text}]codex app-server[/] "
f"[dim {dim}](terminal/file ops/MCP run inside codex)[/]")
# Show active profile name when not 'default'. Never break the banner over a profiles.py bug.
_profile_name = _quiet(_active_profile_name)
if _profile_name and _profile_name != "default":
right_lines.append(f"[bold {accent}]Profile:[/] [{text}]{_profile_name}[/]")
right_lines.append(f"[dim {dim}]{' · '.join(summary_parts)}[/]")
# Update check — NEVER block the banner on it: the prefetch does git/network work that rarely
# finishes before render, so a blocking wait adds its full timeout to every startup. If not
# ready, a daemon thread prints the same notice above the prompt when it lands.
def _update_line():
behind = get_update_result(timeout=0.05)
if behind is None and not _update_check_done.is_set():
_defer_update_notice()
elif behind is not None and behind != 0:
right_lines.append(_format_update_notice(behind))
_quiet(_update_line) # Never break the banner over an update check
layout_table = Table.grid(padding=(0, 2))
layout_table.add_column("left", justify="center")
layout_table.add_column("right", justify="left")
layout_table.add_row("\n".join(left_lines), "\n".join(right_lines))
version_label = format_banner_version_label()
release_info = get_latest_release_tag()
if release_info:
version_label = f"[link={release_info[1]}]{version_label}[/link]"
outer_panel = Panel(
layout_table, title=f"[bold {_skin_color('banner_title', '#FFD700')}]{version_label}[/]",
border_style=_skin_color("banner_border", "#CD7F32"), padding=(0, 2))
console.print()
if shutil.get_terminal_size().columns >= 95:
console.print(getattr(_bskin, "banner_logo", None) or HERMES_AGENT_LOGO)
console.print()
console.print(outer_panel)