Files
hermes-agent/hermes_cli/cli_agent_setup_mixin.py
T
Siddharth Balyan 4bdd64b334 The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract

The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an
anonymous account exactly one model on its own host, refuses everything else with a structured
429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and
tells a signed-in account that still asks for `nous/welcome` what to switch to in an
`x-nous-model-switch` header. Four client-side gaps against that contract:

- Auxiliary calls were refused on every session. The auxiliary client asked the welcome host
  for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free`
  before each fallback. On the welcome host it now uses `nous/welcome` (its backing model
  covers auxiliary work) and skips Nous for vision, which the welcome model does not take.

- The structured 429 body was never read. The classifier now parses `reason` /
  `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are
  non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited`
  are rate limits that honour `retry_after` and never rotate the free tier's only credential.
  The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and
  fall back instead of retrying or re-exchanging. The terminal paths say what happened and
  name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal).

- The `x-nous-model-switch` header was ignored. The chat-completions transport records it
  beside the rate-limit and credits headers; the next call moves the session, and the config
  default when it still names `nous/welcome`, to the backing model the gateway named.

- A guest fell back to the paid host. With `inference_base_url` absent from the exchange or
  outside the host allowlist, routing defaulted to inference-api, where every request is a
  400. A guest now defaults to the welcome literal at the exchange, in the shared store's
  shape, and in effective routing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8)

* feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use

A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided
setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier
picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime
rung now sets it up there instead of failing "not logged in", so the guided chat no longer races
the root profile's first-run mint.

nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever
nothing else is configured; "on-request" mints only when the free tier is asked for by name
(nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers —
the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the
connector token path — still adopt what the shared store holds, so every profile follows the one
identity the guided setup created, but never create one on their own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c)
(cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165)

* feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit

Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity
(nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is
configured. "explicit" means Hermes never creates one on its own: the only creator is the new
provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup
on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and
before the guided chat exists — so the identity lands in the root store every profile reads
through and is there before any session asks for nous/welcome. That closes the race against the
backend's own setup, and makes "only when the setup-bot flow is used" literally true.

The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome
(the hermes model row, --provider nous), which treated a model name as intent and was broader
than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not
logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to
sign in from. Implicit callers still adopt an identity the shared store holds, and a retired
credential is replaced (a continuation, not a creation).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0)

* fix(auth): remove the nous.guest_setup knob; the free tier is created on first use

`nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its
default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`,
unknown values read as `auto`, and under `explicit` a CLI-only install could never get an
identity, which contradicts the first-run contract (first command mints, then chats).

The mint race the knob accompanied is already benign: every caller takes the profile lock then
the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup
win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which
stays. `nous.guest` remains the only free-tier policy.

Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through
`ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call
sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config
entries. The three policy tests that hold regardless of the knob are kept under
`TestExplicitProvision`; the two that only tested the knob are deleted.

(cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261)

* fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent

When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway
serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch`
moves the live session to that model and moves `config.yaml`'s default off the alias in the
same step. The messaging gateway's fallback-eviction check compares the agent's model with the
config default and evicts on any mismatch that is not a /model override, so when the config
write did not land (unreadable config, lock) the cached agent was evicted once per turn, and
prompt caching with it.

`apply_model_switch` now stamps the alias it moved the session off on the agent, and
`_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as
deliberate, beside the existing /model override case. The check takes the agent and the config
model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them.

(cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954)

* fix(auth): the free tier outranks implicit host credentials in provider resolution

On a fresh install with a leftover ~/.aws profile, resolve_provider("auto")
reached the Bedrock rung before the free-tier rung, so the first turn ran on
Bedrock and failed 403 while the free tier was still being minted in the
background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s,
three retries, no answer; the next process then switched to nous/welcome.

The free-tier rung now sits directly above the Bedrock chain: when nous.guest
is on, an existing free-tier identity answers, else a blocking mint runs, and
only then does the boto chain get a say. Everything above is unchanged and
still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter
pool, a logged-in active_provider. nous.guest: false skips the rung, and a
failed mint still falls through to Bedrock and the no-provider guidance.

Tests: six precedence cases (identity present, fresh mint, free tier off, env
key still wins, sign-in still wins, failed mint falls through). The opt-out
test now neutralizes the AWS chain like the precedence tests do; on a machine
with ~/.aws it was failing for the same reason as the bug.

Live after the fix, same Mac, AWS credentials visible, isolated shared store:
identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in
11 s.

(cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a)

* fix(auth): review follow-ups for the free-tier rung (NS-829)

- tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches
  the free tier off; its contract is the boto chain, and the free tier now
  sits above it.
- gateway/run_notifications.py: the free-tier startup line reads auth.json
  before consulting the resolver, so a gateway boot on a machine with AWS
  credentials never mints or refreshes over the network.
- hermes_cli/anon_auth.py: module docstring says where the free tier sits in
  the ladder instead of "the ladder is untouched".
- tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of
  six (parametrized ladder cases; a failed mint that returns None or raises
  falls through to Bedrock).

scripts/run_tests.sh on the five affected files: 147 passed, 0 failed.

(cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11)

* feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone

The free tier is pre-GA. Until GA it must not exist for anyone who did not
ask for it: no identity minted, no portal traffic, no free-tier copy on any
surface. One environment variable now decides that, and one function reads it.

`guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly
"1"; only then does `nous.guest` (the user's off switch) get consulted. Every
free-tier site already funnels through `guest_enabled()`, so the gate closes
minting, routing, connector entitlement, status lines and the picker row in
one place. With the variable unset, `resolve_provider("auto")` on a fresh
install raises `no_provider_configured` exactly as upstream does.

`HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the
gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted
identities as a side effect of provider resolution, and `_has_any_provider_
configured` read them ahead of every other check, making the CLI a second
reader of a flag that must have exactly one. `_forced_new_done` and the
`force` parameter of `_reconcile_and_provision` go with them.

Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in
b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847).

Not a user preference: the variable is never written to config.yaml or .env
and never shown in setup. It is deleted at GA together with its comment in
anon_auth.py. This is a deliberate, temporary exception to the "no new
HERMES_* env vars for non-secret config" rule.

Tests: fixtures set the gate instead of deleting the old lever; one new
invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "",
"0", "true" and "new" all leave the tier off with zero portal calls, red on the
previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the
feature.

* feat(auth): the free-tier identity is created in one place, at boot; every other site is a read

Before this commit eight sites could create a Nous free-tier identity as a
side effect of something else: resolving a provider, the CLI's first-run
check, the CLI's session setup (in the background beside an own key), a
connector bearer read, the desktop polling `free_tier.status`, the sign-in
precondition, the desktop's `free_tier.provision`, and the dead-credential
re-mint. A poll could mint. Provider resolution could hit the network. Two
of them raced each other on a fresh install.

Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator.
`hermes serve` runs it on a daemon thread from `_lifespan` beside the other
background boots; `cmd_chat` runs it synchronously before the first-run
guard. It inventories credentials first (`resolve_provider("auto",
skip_free_tier=True)`: what would carry inference if the free tier did not
exist), creates the identity only when `guest_enabled()`, resolves inference,
records a `SetupRecord` in process memory and broadcasts ONE `setup.ready`
event. It runs on every boot; only the mint is gated.

`ensure_portal_identity` now requires `explicit=True` and raises otherwise.
Its callers are the bootstrap, the desktop's `free_tier.provision` (the
explicit retry when the boot could not create the identity) and the two
dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`,
`managed_tool_gateway._replace_dead_guest_token`). The background thread
path and `provision_free_tier` are deleted with their last callers.

Reads that used to mint and now only read: `auth.py::resolve_provider`
rung 7 (an existing identity still outranks the Bedrock chain, NS-829
ordering kept), `main.py::_has_any_provider_configured`,
`cli_agent_setup_mixin._ensure_runtime_credentials`,
`managed_tool_gateway.read_nous_access_token` (no identity -> None),
`anon_sign_in.run_sign_in` (no identity -> Unavailable),
`methods_free_tier` `free_tier.status`.

`setup.status` answers from the record for the launch profile, blocking up
to 8 s while the bootstrap is in flight so a client's first poll lands after
the identity exists rather than racing it; a named profile, or a process
that never ran the bootstrap, keeps today's live probe. The record's fields
ride along additively (`ready`, `free_tier`, `other_providers`,
`inference_provider`).

Identity and inference are decoupled (NS-845 Q1.3): the mint sets
`active_provider="nous"` only when the inventory found nothing else usable
(`_mint_locked(carries_inference=)`); an adopted account always does. A token
refresh no longer re-elects the provider it refreshed
(`_save_provider_state_to_source` writes credentials, not the user's
choice) — that write was how an own-key install ended up on the free tier
after the first connector call.

Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check),
bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint),
62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and
`provision_free_tier`), and a04b05260c (blocking mint in the resolver).
Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847.

Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps
inference; reads never reach the portal; a refused mint is memoised),
`free_tier.status` fails loudly if it ever calls the creator, the resolver
stub fails loudly if resolution ever mints, `setup.status` reads the record,
`skip_free_tier` proves the inventory question. The three sign-in tests for
the deleted pre-mint collapse into one (`no identity -> Unavailable, zero
portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and
off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI
matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20.

* fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup"

A free-tier identity carries $0 by design, so the portal seed reports
`paid_access=False` for it. `is_free_tier_model` did not know the welcome
host, read that as a depleted account, and every free-tier turn ended with
the credits-depleted notice telling the user to top up an account they do
not have.

Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host
(`anon_auth.route_is_welcome_host`) is the free tier. The host is the
evidence, not the model name: the paid inference host can serve
`nous/welcome` to a named account and that account's depletion is real, so
`("nous/welcome", <inference host>)` stays False. Local data only, like the
three rules above it.

Restores the two contracts dropped by hermes-magic 674e11d1eaa (the
prototype line ran without unit tests): the welcome host is free without
any pricing evidence; the model name alone is not. The first is red without
this fix.

* fix(copy): free-tier text stops promising a connector transfer and never names the config key

Sign-in copy on every surface said "Sign in to keep your connectors" and
ended with "Your connectors are kept." The transfer registry that would
make that true is empty (NS-821): nothing carries over today. The copy now
says what signing in does give ("unlock more models and tools") and the
completion line names the account, not a transfer. The docs page loses the
"connectors carry over" paragraph for the same reason.

The picker's off-state line exposed `nous.guest: false` and the word
"guest"; user copy names the free tier only (R-USR-1).

The docs page gains the pre-rollout note: until GA nothing on it happens
without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and
"replaced on next use" sentences now describe the boot bootstrap.

zh is a strict locale: the `freeTier` block was English placeholder text
copied from `en`; it is now Chinese. `connectorsKept` is renamed
`completedBody` since it no longer talks about connectors.

* feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn

The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1"
as on. Until now nothing in the desktop set it, so a packaged app could
never turn the free tier on, and a backend spawned by the app could
disagree with the app about whether the tier was live.

`electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled`
is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has
`--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch
into a module constant. `desktopBackendSpawnEnv` wraps every backend env
as the outermost call and writes the flag LAST, as "1" or an explicit "0",
so no earlier spread (`process.env`, `backend.env`) can resurrect a stray
value from the parent shell.

Stamped onto all three spawn sites: the primary `serve` spawn, the pooled
per-profile spawn, and the remote SSH `exec env ...` command (which gains
` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and
the backend probes are not backend spawns and do not get it: a
`hermes --tui` typed in the pane must not mint.

The renderer learns the same fact read-only through the existing
`hermes:launch-flags` sync IPC (`guestOnboarding`) and preload
(`window.hermesDesktop.guestOnboardingEnabled`).

Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding`
maps to it in main). Two invariant tests on the pure helpers: only "1" or
the argv flag enables; the spawn env carries "1"/"0" as the last word and
preserves every other key.

* feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll

The backend's boot bootstrap now announces `setup.ready` once, after it has
created (or refused) the free-tier identity and resolved the inference
route. The renderer used to discover both by polling `setup.status`,
`setup.runtime_check` and `free_tier.status` every 60 s from
`useStatusSnapshot`; a fresh install's chip, notice strip and onboarding
overlay could sit stale for up to a minute after boot, and three RPCs a
minute per window kept asking a question whose answer changes only at
boundaries the backend already announces.

`handleLifecycleEvent` routes `setup.ready` (active source only, like
`skin.changed`) to `notifySetupReady()`, a one-shot tick atom in
`live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens
to it and runs one readiness round at once (`setup.status` +
`setup.runtime_check` + `free_tier.status`). The readiness legs also run
once on open and on return from another app, as today. The 60 s tick keeps
only `getStatus()`.

`SetupStatusSnapshot` types the record's additive fields (`ready`,
`free_tier`, `other_providers`, `inference_provider`); readiness semantics
are unchanged and still key on `provider_configured` + `runtime_check`.

Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one
refresh from the active source and none from another; the snapshot hook's
contract is three legs on open, one leg on the tick.

* fix(cli): the banner names the free tier's model instead of "no model configured"

The welcome banner prints before credentials resolve, so on a fresh install
`model` is empty and the banner said, in red, "no model configured — run
/model or hermes setup". Under the free tier that is false: the route is
already known from local state (identity on disk, tier on), and the first
message will run on `nous/welcome`.

`_banner_left_lines` now asks the route the same question when `model` is
empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous
Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the
banner's 'no model configured' line reads the resolved route").

Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`;
gate off -> the red line, zero portal calls.

* fix(aux): vision on the free tier uses nous/welcome too

The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the
backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the
repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the
welcome host would have sent every image step past the free tier for no reason, so the
auxiliary client pins the route's one model for every lane. A backing model that takes
no images answers with the upstream's own error, which the ladder handles as it always has.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415)

* fix(gateway): hermes gateway run is a boot owner of the free tier too

Rung 5 made every demand-time free-tier site a read: resolve_provider,
the connector token, the /login precondition. That is only correct if
every process that can reach those sites ran the bootstrap first. The
CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone
messaging gateway did not. A fresh HERMES_HOME with the gate on and
`hermes gateway run` reached provider resolution with no identity to
consume, and /login returned Unavailable. Reported by @andrexibiza on
#107697 (P1).

GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an
executor thread right after startup recovery and BEFORE any adapter
connects, so a fast first DM cannot arrive with nothing to resolve. It
is its own step, not part of the turn-machinery warm-up: the warm-up is
an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0);
the bootstrap is correctness and must always run. With the gate unset it
is a local inventory and no network.

Live, real GatewayRunner.start against a fake portal in a fresh home:
  gate on   -> 1 create, identity persisted, resolve_runtime_provider=nous,
               /login precondition sees the identity
  gate off  -> 0 portal calls, no identity, no_provider_configured
Before the fix the gate-on row was identical to the gate-off row.

Test: the bootstrap seam runs before _start_prefilter_platforms and
delegates to the one creator. Red on 5554eb6993 (no seam), green here.

---------

Co-authored-by: Robin Fernandes <robin@soal.org>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 03:45:33 +05:30

733 lines
38 KiB
Python

"""Agent construction + session-resume display for ``HermesCLI``: credential resolution,
per-turn agent config, first-use build, resume preload + recap. ``cli.py`` helpers are
imported lazily inside each method (import cycle)."""
from __future__ import annotations
import sys
from rich.markup import escape as _escape
from utils import base_url_host_matches
def _single_query_clarify_callback(question: str, choices=None, multi_select=False) -> str:
"""Headless clarify answer for ``hermes chat -q``.
A -q turn never builds the prompt_toolkit app, so the interactive clarify modal
can never be painted or answered — the CLI callback would poll until
``agent.clarify_timeout`` while the caller sees a silent hang. Mirror the oneshot
path and answer immediately instead.
The oneshot path answers immediately via ``_oneshot_clarify_callback``; single-query turns need the same
headless behavior (#94943).
"""
prefix = f"[single-query mode: no user available to answer {question!r}. "
if choices:
what = "subset" if multi_select else "option"
return f"{prefix}Pick the best {what} from {choices} using your own judgment and continue.]"
return f"{prefix}Make the most reasonable assumption you can and continue.]"
def _current_runtime(cli) -> dict:
"""Snapshot the CLI's resolved provider routing as an AIAgent runtime dict.
getattr guards stay: tests build minimal shells lacking these attributes."""
return {
"api_key": cli.api_key,
"base_url": cli.base_url,
"provider": cli.provider,
"requested_provider": getattr(cli, "requested_provider", cli.provider),
"api_mode": cli.api_mode,
"command": cli.acp_command,
"args": list(cli.acp_args or []),
"credential_pool": getattr(cli, "_credential_pool", None)}
def _route_signature(model, runtime: dict) -> tuple:
"""Hashable identity of (model, routing) used to detect when the agent must be rebuilt."""
return (
model, runtime.get("provider"), runtime.get("requested_provider"), runtime.get("base_url"),
runtime.get("api_mode"), runtime.get("command"), tuple(runtime.get("args") or ()))
def _keyless_custom_base(base_url) -> bool:
"""Custom/local endpoints (llama.cpp, ollama, vLLM) often need no auth; only a
non-OpenRouter base_url qualifies."""
return bool(
isinstance(base_url, str)
and base_url
and not base_url_host_matches(base_url, "openrouter.ai"))
def _compression_descendant(session_db, session_id):
"""If ``session_id`` is the (empty) head of a compression chain, return the
descendant that actually holds the messages; else None. Fails open on DB errors."""
try:
resolved_id = session_db.resolve_resume_session_id(session_id)
except Exception:
return None
return resolved_id if resolved_id and resolved_id != session_id else None
def _user_display_text(content) -> str:
"""Recap text for a user row; multimodal lists become text parts + ``[image]`` markers."""
if isinstance(content, list):
return " ".join(
part.get("text", "") if part.get("type") == "text" else "[image]"
for part in content
if isinstance(part, dict) and part.get("type") in ("text", "image_url"))
return "" if content is None else str(content)
def _tool_calls_summary(tool_calls) -> str:
"""``[N tool call(s): name, ...]`` with up to 4 distinct names."""
names = []
for tc in tool_calls:
fn = tc.get("function", {})
name = fn.get("name", "unknown") if isinstance(fn, dict) else "unknown"
if name not in names:
names.append(name)
names_str = ", ".join(names[:4]) + (", ..." if len(names) > 4 else "")
noun = "call" if len(tool_calls) == 1 else "calls"
return f"[{len(tool_calls)} tool {noun}: {names_str}]"
# display_kind -> recap event line; ``hidden`` rows are skipped before this lookup.
_RESUME_EVENT_TEXT = {
"model_switch": "model changed",
"async_delegation_complete": "background delegation completed",
"auto_continue": "resumed interrupted turn"}
def _collect_resume_entries(display_history, disp: dict, clean_assistant):
"""Displayable ``(role, text)`` recap entries from stored history, truncated per the
``display.resume_*`` config; system and tool-result rows are skipped. Returns
``(entries, index of last assistant entry, its un-truncated text)``.
Stored history is untrusted for display: text is sanitized so replay can't clear the
screen, retitle the window or restyle the panel. Pure-reasoning assistant rows with no
visible output are skipped, as are tool-call-only rows when ``resume_skip_tool_only``.
"""
from tools.ansi_strip import sanitize_display_text as _sanitize_display_text
max_user_len = int(disp.get("resume_max_user_chars", 300))
max_asst_len = int(disp.get("resume_max_assistant_chars", 200))
max_asst_lines = int(disp.get("resume_max_assistant_lines", 3))
skip_tool_only = disp.get("resume_skip_tool_only", True)
entries: list = []
last_asst_idx = None
last_asst_full = None
for msg in display_history:
role = msg.get("role", "")
display_kind = msg.get("display_kind")
content = msg.get("content")
tool_calls = msg.get("tool_calls") or []
if display_kind == "hidden":
continue
if display_kind in _RESUME_EVENT_TEXT:
metadata = msg.get("display_metadata") or {}
label = metadata.get("display_text") if display_kind == "async_delegation_complete" else None
entries.append(("event", _sanitize_display_text(label or _RESUME_EVENT_TEXT[display_kind])))
continue
if role == "user":
text = _sanitize_display_text(_user_display_text(content))
if len(text) > max_user_len:
text = text[:max_user_len] + "..."
entries.append(("user", text))
elif role == "assistant":
text = clean_assistant("" if content is None else str(content))
parts, full_parts = [], []
if text:
full_parts.append(text)
lines = text.splitlines()
if len(lines) > max_asst_lines:
text = "\n".join(lines[:max_asst_lines]) + " ..."
if len(text) > max_asst_len:
text = text[:max_asst_len] + "..."
parts.append(text)
if tool_calls:
parts.append(_tool_calls_summary(tool_calls))
full_parts.append(parts[-1])
if not text and (skip_tool_only or not tool_calls):
continue
entries.append(("assistant", " ".join(parts)))
last_asst_idx = len(entries) - 1
last_asst_full = " ".join(full_parts)
return entries, last_asst_idx, last_asst_full
# (skin key, fallback) for recap panel colors: body text, session label, border, assistant label.
_RESUME_SKIN_COLORS = (
("banner_text", "#FFF8DC"), ("session_label", "#DAA520"), ("session_border", "#8B8682"),
("ui_ok", "#8FBC8F"))
def _resume_panel_colors() -> tuple:
"""Active-skin colors for ``_RESUME_SKIN_COLORS`` (fallbacks when no skin loads)."""
try:
from hermes_cli.skin_engine import get_active_skin
_skin = get_active_skin()
return tuple(_skin.get_color(key, default) for key, default in _RESUME_SKIN_COLORS)
except Exception:
return tuple(default for _, default in _RESUME_SKIN_COLORS)
class CLIAgentSetupMixin:
"""Agent construction + session-resume display methods for ``HermesCLI``."""
def _ensure_runtime_credentials(self) -> bool:
"""Re-resolve provider credentials before agent use so key rotation / token
refresh are picked up without restarting the CLI. False on auth failure."""
from cli import ChatConsole, logger
from hermes_cli.runtime_provider import resolve_runtime_provider, format_runtime_provider_error
_primary_exc = None
runtime = None
try:
runtime = resolve_runtime_provider(
requested=self.requested_provider, explicit_api_key=self._explicit_api_key,
explicit_base_url=self._explicit_base_url)
except Exception as exc:
_primary_exc = exc
if _primary_exc is not None:
runtime = self._resolve_fallback_runtime(_primary_exc)
if runtime is not None:
_primary_exc = None
if runtime is None:
message = format_runtime_provider_error(_primary_exc) if _primary_exc else "Provider resolution failed."
ChatConsole().print(f"[bold red]{message}[/]")
return False
api_key = runtime.get("api_key")
base_url = runtime.get("base_url")
resolved_provider = runtime.get("provider", "openrouter")
if resolved_provider != "nous":
# An explicit provider carries inference. The free-tier identity (for connectors) was
# created by the boot bootstrap before this point, never here; this prints the one-time
# "free tier is here" notice the first time an identity is seen beside an own key.
self._maybe_print_free_tier_available_notice()
resolved_routing = (
resolved_provider, runtime.get("api_mode", self.api_mode), runtime.get("command"),
list(runtime.get("args") or []))
# A callable api_key is a bearer-token provider (Azure Entra ID): the OpenAI SDK
# invokes it per request, so skip string validation / placeholder substitution.
if not callable(api_key) and not (isinstance(api_key, str) and api_key):
if _keyless_custom_base(base_url):
# Placeholder key so the SDK doesn't reject the keyless local endpoint.
api_key = "no-key-required"
logger.debug(
"No API key for custom endpoint %s (source=%s), "
"using placeholder — local servers typically ignore auth",
base_url, runtime.get("source", ""))
else:
_prov = (resolved_provider or self.requested_provider or "").strip()
if _prov and _prov != "auto":
print(f"\n⚠️ No API key found for provider '{_prov}'.")
else:
print("\n⚠️ No inference provider is configured.")
print(" Run 'hermes model' to choose a provider, or "
"'hermes setup' for first-time setup.")
return False
if not isinstance(base_url, str) or not base_url:
print("\n⚠️ Provider resolver returned an empty base URL. "
"Check your provider config or run: hermes setup")
return False
credentials_changed = api_key != self.api_key or base_url != self.base_url
routing_changed = resolved_routing != (self.provider, self.api_mode, self.acp_command, self.acp_args)
self.provider, self.api_mode, self.acp_command, self.acp_args = resolved_routing
self._credential_pool = runtime.get("credential_pool")
self._provider_source = runtime.get("source")
self.api_key = api_key
self.base_url = base_url
# A custom_provider entry's explicit `model` wins when the CLI model is unset or
# is just the provider slug/display name (`hermes chat --model <provider-name>`
# would otherwise send the provider name as the model string -> 400).
runtime_model = runtime.get("model")
if runtime_model and isinstance(runtime_model, str) and (
not self.model or self.model == self.provider or self.model == runtime.get("name")):
self.model = runtime_model
# Still empty (e.g. `hermes auth add` without `hermes model`): fall back to the
# provider's first catalog model so the API doesn't reject an empty model.
if not self.model and resolved_provider:
try:
from hermes_cli.models import get_default_model_for_provider
_default = get_default_model_for_provider(resolved_provider)
if _default:
self.model = _default
logger.info(
"No model configured — defaulting to %s for provider %s",
_default, resolved_provider)
except Exception:
pass
# Normalize model for the resolved provider (e.g. swap non-Codex models on openai-codex).
# Fixes #651.
model_changed = self._normalize_model_for_provider(resolved_provider)
# AIAgent/OpenAI client holds auth at init, so rebuild on key/routing/model change.
if (credentials_changed or routing_changed or model_changed) and self.agent is not None:
self.agent = None
self._active_agent_route_signature = None
return True
def _maybe_print_free_tier_available_notice(self) -> None:
"""One-time notice for installs whose inference is carried by an explicit provider: the free
tier (inference + connectors) now exists. Printed the first time an identity is present, then
flagged on that identity so it never repeats. Never blocks or raises."""
from cli import logger
try:
from hermes_cli import anon_auth
if not anon_auth.guest_notice_pending():
return
self._console_print(f"[dim]{anon_auth.FREE_TIER_AVAILABLE_NOTICE}[/]")
anon_auth.mark_guest_notice_shown()
except Exception as exc:
logger.debug("free tier availability notice skipped: %s", exc)
def _resolve_fallback_runtime(self, primary_exc):
"""Primary provider resolution failed: on an AuthError try each fallback entry in
order and switch the CLI's requested_provider/model to the first that resolves.
None when the error is not auth-related or no fallback resolves."""
from cli import _cprint, logger
from hermes_cli.auth import AuthError
from hermes_cli.runtime_provider import resolve_runtime_provider
if not isinstance(primary_exc, AuthError):
return None
_fb_chain = self._fallback_model if isinstance(self._fallback_model, list) else []
for _fb in _fb_chain:
_fb_provider = (_fb.get("provider") or "").strip().lower()
_fb_model = (_fb.get("model") or "").strip()
if not _fb_provider or not _fb_model:
continue
try:
from hermes_cli.fallback_config import resolve_entry_api_key
_fb_kwargs = {"requested": _fb_provider}
if _fb.get("base_url"):
_fb_kwargs["explicit_base_url"] = _fb["base_url"]
_fb_api_key = resolve_entry_api_key(_fb)
if _fb_api_key:
_fb_kwargs["explicit_api_key"] = _fb_api_key
runtime = resolve_runtime_provider(**_fb_kwargs)
logger.warning(
"Primary provider auth failed (%s). Falling through to fallback: %s/%s",
primary_exc, _fb_provider, _fb_model)
_cprint(f"⚠️ Primary auth failed — switching to fallback: {_fb_provider} / {_fb_model}")
self.requested_provider = _fb_provider
self.model = _fb_model
return runtime
except Exception:
continue
return None
def _runtime_credentials_ready(self) -> bool:
"""Silently probe whether any inference provider can be resolved.
Never prints or mutates CLI state, so the interactive first-run path can route a
keyless install into onboarding before the user types into a chat that can't work.
See #62935.
"""
from hermes_cli.runtime_provider import resolve_runtime_provider
try:
runtime = resolve_runtime_provider(
requested=self.requested_provider, explicit_api_key=self._explicit_api_key,
explicit_base_url=self._explicit_base_url)
except Exception:
return False
if not isinstance(runtime, dict):
return False
api_key = runtime.get("api_key")
base_url = runtime.get("base_url")
if callable(api_key) or (isinstance(api_key, str) and api_key):
return bool(base_url)
return _keyless_custom_base(base_url)
def _offer_first_run_setup(self) -> bool:
"""Offer the provider picker when no provider is configured at all (interactive
startup, TTY). Runs the same flow as ``hermes model`` so onboarding has a single
source of truth. True when a provider was configured."""
from cli import _cprint, logger
_cprint("")
_cprint("⚕ No inference provider is configured yet — let's fix that.")
_cprint(" You'll pick a provider (Nous Portal OAuth is the fastest; "
"no API key needed) and a model.")
try:
answer = input(" Set up a provider now? [Y/n]: ").strip().lower()
except (KeyboardInterrupt, EOFError):
print()
answer = "n"
if answer in {"n", "no"}:
_cprint(" Skipped. Run 'hermes model' or 'hermes setup' any time.")
return False
try:
from hermes_cli.main import select_provider_and_model
select_provider_and_model()
except (KeyboardInterrupt, EOFError, SystemExit):
print()
_cprint(" Setup cancelled. Run 'hermes model' any time.")
return False
except Exception as exc:
logger.debug("first-run provider setup failed: %s", exc)
_cprint(f" ⚠️ Provider setup failed: {exc}")
_cprint(" Run 'hermes model' to try again.")
return False
# Re-sync CLI state from what the picker persisted so the next turn uses it without a restart.
try:
from hermes_cli.config import load_config
_model_cfg = (load_config().get("model") or {})
if isinstance(_model_cfg, dict):
self.requested_provider = (_model_cfg.get("provider") or "").strip() or self.requested_provider
_new_model = (_model_cfg.get("default") or _model_cfg.get("model") or "").strip()
self.model = _new_model or self.model
except Exception as exc:
logger.debug("first-run config re-sync failed: %s", exc)
# Force credential re-resolution + agent rebuild on next use.
self.agent = None
self._active_agent_route_signature = None
if self._runtime_credentials_ready():
_cprint(" ✓ Provider configured — you're ready to chat.")
return True
_cprint(" Provider setup didn't complete. Run 'hermes model' to retry.")
return False
def _resolve_turn_agent_config(self, user_message: str) -> dict:
"""Effective model/runtime config for one turn — always the session's primary
provider. With `/fast` on (service_tier == "priority") attach request_overrides;
auto/cold tiers are applied per request by agent.fast_mode instead."""
from hermes_cli.models import resolve_fast_mode_overrides
runtime = _current_runtime(self)
route = {"model": self.model, "runtime": runtime, "signature": _route_signature(self.model, runtime)}
overrides = None
if getattr(self, "service_tier", None) == "priority":
try:
overrides = resolve_fast_mode_overrides(
route["model"], provider=runtime["provider"], base_url=runtime["base_url"])
except Exception:
pass
route["request_overrides"] = overrides
return route
def _follow_compression_chain(self, session_meta, announce):
"""If the resumed id is an empty compression-chain head, announce and switch to
the descendant holding the messages; returns the (possibly refreshed) meta."""
resolved_id = _compression_descendant(self._session_db, self.session_id)
if resolved_id:
announce(resolved_id)
self.session_id = resolved_id
session_meta = self._session_db.get_session(self.session_id) or session_meta
return session_meta
def _restore_session_state(self, session_meta, *, quiet: bool = False) -> None:
"""Restore cwd / yolo / model from the resumed session's metadata."""
self._restore_session_cwd(session_meta, quiet=quiet)
self._restore_session_yolo(session_meta, quiet=quiet)
self._restore_session_model(session_meta, quiet=quiet)
def _reopen_session(self) -> None:
"""Clear ended_at so the resumed session is active again (best effort)."""
try:
self._session_db.reopen_session(self.session_id)
except Exception:
pass
def _load_resumed_history_late(self) -> bool:
"""Late resume path: validate the session and load its history from the DB when
_preload_resumed_session() (called from run()) did not already populate it.
False when the resume must abort (missing session / over the safe-resume limit)."""
from cli import ChatConsole, _DIM, _RST, _accent_hex, _cprint
session_meta = self._session_db.get_session(self.session_id)
# Quiet mode (tool_progress_mode == "off") routes resume status lines to
# stderr so stdout stays machine-readable for `$(hermes chat -Q --resume ...)`.
# Without this, the resume banner pollutes captured stdout. See #11793.
_quiet_mode = getattr(self, "tool_progress_mode", "full") == "off"
def _say(plain: str, rich: str) -> None:
if _quiet_mode:
print(plain, file=sys.stderr)
else:
ChatConsole().print(rich)
if not session_meta:
hint = "Use a session ID from a previous CLI run (hermes sessions list)."
if _quiet_mode:
print(f"Session not found: {self.session_id}", file=sys.stderr)
print(hint, file=sys.stderr)
else:
_cprint(f"\033[1;31mSession not found: {self.session_id}{_RST}")
_cprint(f"{_DIM}{hint}{_RST}")
return False
session_meta = self._follow_compression_chain(
session_meta,
lambda rid: ChatConsole().print(
f"[dim]Session {_escape(self.session_id)} was compressed into "
f"{_escape(rid)}; resuming the descendant with your "
f"transcript.[/dim]"))
if getattr(self, "_resume_history_error", None):
return False
# Only the TIP session's rows are loaded here (no ancestors), so use the
# tip-only count — the full-lineage count would over-reject compressed sessions.
resume_limit_error = self._resume_history_limit_error(tip_only=True)
if resume_limit_error:
self._resume_history_error = resume_limit_error
_say(
f"Cannot resume session: {resume_limit_error}",
f"[bold red]Cannot resume session:[/] {_escape(resume_limit_error)}")
return False
restored = self._session_db.get_messages_as_conversation(self.session_id, repair_alternation=True)
if restored:
restored = [m for m in restored if m.get("role") != "session_meta"]
self.conversation_history = restored
msg_count = len([m for m in restored if m.get("role") == "user"])
title_part = f" \"{session_meta['title']}\"" if session_meta.get("title") else ""
counts = f"({msg_count} user message{'s' if msg_count != 1 else ''}, {len(restored)} total messages)"
_say(
f"↻ Resumed session {self.session_id}{title_part} {counts}",
f"[bold {_accent_hex()}]↻ Resumed session[/] [bold]{_escape(self.session_id)}[/]"
f"[bold {_accent_hex()}]{_escape(title_part)}[/] {counts}")
self._restore_session_state(session_meta, quiet=_quiet_mode)
else:
_say(
f"Session {self.session_id} found but has no messages. Starting fresh.",
f"[bold {_accent_hex()}]Session {_escape(self.session_id)} found but has no messages. Starting fresh.[/]",
)
self._reopen_session()
return True
def _init_agent(self, *, model_override: str = None, runtime_override: dict = None, request_overrides: dict | None = None) -> bool:
"""Build the agent on first use; when resuming, restore history from SQLite.
Returns True on success."""
from cli import ChatConsole, _cprint, _prepare_deferred_agent_startup, logger
from run_agent import AIAgent
if self.agent is not None:
return True
# Join the background preloaded-skills load (--skills/-s) BEFORE the agent
# snapshots self.system_prompt below. No-op when nothing was requested.
self.finalize_preloaded_skills()
_prepare_deferred_agent_startup()
self._install_tool_callbacks()
self._ensure_tirith_security()
if not self._ensure_runtime_credentials():
return False
from hermes_cli.mcp_startup import ensure_mcp_discovery_before_agent_build
ensure_mcp_discovery_before_agent_build(
logger=logger, single_query=getattr(self, "_single_query_mode", False))
if self._session_db is None:
try:
from hermes_state import SessionDB
self._session_db = SessionDB()
except Exception as e:
logger.warning("SQLite session store not available — session will NOT be indexed: %s", e)
if (
self._resumed and self._session_db and not self.conversation_history
and not self._load_resumed_history_late()):
return False
try:
runtime = runtime_override or _current_runtime(self)
effective_model = model_override or self.model
# -q never builds the prompt_toolkit app, so the clarify modal can't be
# answered — answer headless instead of polling until clarify_timeout.
clarify_callback = (
# See #94943.
_single_query_clarify_callback
if getattr(self, "_single_query_mode", False)
else self._clarify_callback)
self.agent = AIAgent(
model=effective_model, api_key=runtime.get("api_key"),
base_url=runtime.get("base_url"), provider=runtime.get("provider"),
requested_provider=runtime.get("requested_provider"),
api_mode=runtime.get("api_mode"), acp_command=runtime.get("command"),
acp_args=runtime.get("args"), credential_pool=runtime.get("credential_pool"),
max_iterations=self.max_turns,
run_budget_seconds=getattr(self, "run_budget_seconds", None),
enabled_toolsets=self.enabled_toolsets, disabled_toolsets=self.disabled_toolsets,
verbose_logging=self.verbose, quiet_mode=not self.verbose,
tool_progress_mode=getattr(self, "tool_progress_mode", "all"),
ephemeral_system_prompt=self.system_prompt if self.system_prompt else None,
prefill_messages=self.prefill_messages or None,
reasoning_config=self.reasoning_config, service_tier=self.service_tier,
request_overrides=request_overrides, providers_allowed=self._providers_only,
providers_ignored=self._providers_ignore, providers_order=self._providers_order,
provider_sort=self._provider_sort,
provider_require_parameters=self._provider_require_params,
provider_data_collection=self._provider_data_collection,
openrouter_min_coding_score=self._openrouter_min_coding_score,
session_id=self.session_id, platform="cli", session_db=self._session_db,
clarify_callback=clarify_callback,
reasoning_callback=self._current_reasoning_callback(),
fallback_model=self._fallback_model, thinking_callback=self._on_thinking,
checkpoints_enabled=self.checkpoints_enabled,
checkpoint_max_snapshots=self.checkpoint_max_snapshots,
checkpoint_max_total_size_mb=self.checkpoint_max_total_size_mb,
checkpoint_max_file_size_mb=self.checkpoint_max_file_size_mb,
pass_session_id=self.pass_session_id, skip_context_files=self.ignore_rules,
skip_memory=self.ignore_rules, tool_progress_callback=self._on_tool_progress,
tool_start_callback=self._on_tool_start if self._inline_diffs_enabled else None,
tool_complete_callback=self._on_tool_complete if self._inline_diffs_enabled else None,
stream_delta_callback=self._stream_delta if self.streaming_enabled else None,
tool_gen_callback=self._on_tool_gen_start if self.streaming_enabled else None,
notice_callback=self._on_notice, notice_clear_callback=self._on_notice_clear,
reaction_callback=self._on_reaction)
# Reference for atexit memory-provider shutdown: ``_run_cleanup`` in cli.py
# reads ``cli._active_agent_ref``, so this MUST write the ``cli`` module's
# global — a ``global`` statement here would bind this module's namespace.
# When this code lived in cli.py a bare ``global _active_agent_ref`` worked; after the god-file
# extraction into this mixin a ``global`` here would bind *this module's* namespace, leaving
# ``cli._active_agent_ref`` None forever — so memory shutdown never ran on /exit (#49287).
import cli as _cli
_cli._active_agent_ref = self.agent
# Route agent status output through prompt_toolkit so ANSI escapes aren't garbled by
# patch_stdout's StdoutProxy (#2262), holding lines while a response box streams so a
# subagent/background completion notice never splits the reply mid-paragraph.
self.agent._print_fn = self._agent_status_print
# Hydrate credits notices at session OPEN (parity with the TUI) so a depletion
# warning shows before the first message. Idempotent + fail-open in the helper.
try:
from agent.credits_tracker import seed_credits_at_session_start
seed_credits_at_session_start(self.agent)
except Exception:
pass
self._active_agent_route_signature = _route_signature(effective_model, runtime)
# Force-create DB row on /title intent, then apply title.
if self._pending_title and self._session_db:
try:
self.agent._ensure_db_session()
if self.agent._session_db_created:
self._session_db.set_session_title(self.session_id, self._pending_title)
_cprint(f" Session title applied: {self._pending_title}")
self._pending_title = None
# else: row creation failed transiently — keep _pending_title for retry
except Exception as e:
_cprint(f" Could not apply pending title: {e}")
# Keep _pending_title so it can be retried after row creation succeeds
return True
except Exception as e:
console = ChatConsole()
console.print(f"[bold red]Failed to initialize agent: {e}[/]")
from hermes_constants import partial_update_hint
for line in partial_update_hint(e):
console.print(line)
return False
def _resume_history_limit_error(self, tip_only: bool = False):
"""Return a safe-resume error without materializing transcript rows.
``tip_only`` matches call sites that load only the tip session's rows — counting
the full lineage there would over-reject heavily-compressed sessions with a small
tip. Generic guard failures fail OPEN; only a genuine over-limit result blocks."""
if not self._session_db:
return None
from cli import logger
from hermes_state import SessionResumeTooLargeError
try:
safety_check = getattr(self._session_db, "assert_resume_safe", None)
if not callable(safety_check):
return None
safety_check(self.session_id, **({"tip_only": True} if tip_only else {}))
except SessionResumeTooLargeError as exc:
return str(exc)
except Exception as exc:
logger.warning(
"Resume safety check failed for %s (proceeding without guard): %s",
self.session_id, exc)
return None
def _preload_resumed_session(self) -> bool:
"""Load a resumed session's history early (from run(), before the first chat) so
it can be displayed; ``_init_agent()`` then skips its own DB round-trip. Sets
``self.conversation_history`` and prints the status line. True if history loaded."""
from cli import _accent_hex
if not self._resumed or not self._session_db:
return False
session_meta = self._session_db.get_session(self.session_id)
if not session_meta:
self._console_print(f"[bold red]Session not found: {self.session_id}[/]")
self._console_print("[dim]Use a session ID from a previous CLI run (hermes sessions list).[/]")
return False
session_meta = self._follow_compression_chain(
session_meta,
lambda rid: self._console_print(
f"[dim]Session {self.session_id} was compressed into "
f"{rid}; resuming the descendant with your transcript.[/]"))
resume_limit_error = self._resume_history_limit_error()
if resume_limit_error:
self._resume_history_error = resume_limit_error
self._console_print(f"[bold red]Cannot resume session:[/] {resume_limit_error}")
return False
restored, display_history = self._session_db.get_resume_conversations(self.session_id)
accent_color = _accent_hex()
if not restored:
self._console_print(
f"[{accent_color}]Session {self.session_id} found but has no "
f"messages. Starting fresh.[/]")
return False
restored = [m for m in restored if m.get("role") != "session_meta"]
self.conversation_history = restored
self._resume_display_history = [m for m in display_history if m.get("role") != "session_meta"]
from agent.context_compressor import is_user_originated_turn
# Count only user-originated turns: legacy compaction handoffs are durable
# role=user rows without display_kind.
msg_count = len([m for m in self._resume_display_history if is_user_originated_turn(m)])
title_part = f' "{session_meta["title"]}"' if session_meta.get("title") else ""
self._console_print(
f"[{accent_color}]↻ Resumed session [bold]{self.session_id}[/bold]"
f"{title_part} "
f"({msg_count} user message{'s' if msg_count != 1 else ''}, "
f"{len(restored)} total messages)[/]")
self._restore_session_state(session_meta)
self._reopen_session()
return True
def _display_resumed_history(self):
"""Render a dim Rich-panel recap of the previous conversation, capped at the last
``resume_exchanges`` user/assistant exchanges with a hidden-count indicator."""
from cli import CLI_CONFIG, _record_output_history_entry, _strip_reasoning_tags, _suspend_output_history
from tools.ansi_strip import sanitize_display_text as _sanitize_display_text
display_history = getattr(self, "_resume_display_history", self.conversation_history)
if not display_history or self.resume_display == "minimal":
return
_disp = CLI_CONFIG.get("display", {})
entries, _last_asst_idx, _last_asst_full = _collect_resume_entries(
display_history, _disp, lambda t: _sanitize_display_text(_strip_reasoning_tags(t)))
if not entries:
return
skipped = max(0, len(entries) - int(_disp.get("resume_exchanges", 10)) * 2)
entries = entries[skipped:]
# Show the last assistant entry in full so the user sees where they left off.
if _last_asst_idx is not None and _last_asst_full:
adj_idx = _last_asst_idx - skipped
if 0 <= adj_idx < len(entries):
entries[adj_idx] = ("assistant_last", _last_asst_full)
from rich.panel import Panel
from rich.text import Text
_history_text_c, _session_label_c, _session_border_c, _assistant_label_c = (
_resume_panel_colors())
# role -> (label, label style, body style, continuation indent)
role_styles = {
"user": (" ● You: ", f"dim bold {_session_label_c}", "dim", " " * 9),
"assistant": (" ◆ Hermes: ", f"dim bold {_assistant_label_c}", "dim", " " * 12),
"assistant_last": (" ◆ Hermes: ", f"bold {_assistant_label_c}", "", " " * 12), # full, non-dim
}
lines = Text()
if skipped:
lines.append(f" ... {skipped} earlier messages ...\n\n", style="dim italic")
for i, (role, text) in enumerate(entries):
if role == "event":
lines.append(f" ◈ {text}\n", style="dim italic")
else:
label, label_style, body_style, indent = role_styles[role]
lines.append(label, style=label_style)
first, *rest = text.splitlines() or [""] # first line inline, rest indented
lines.append(first + "\n", style=body_style)
for ml in rest:
lines.append(f"{indent}{ml}\n", style=body_style)
if i < len(entries) - 1:
lines.append("") # small gap
panel = Panel(
lines, title=f"[dim {_session_label_c}]Previous Conversation[/]",
border_style=f"dim {_session_border_c}", padding=(0, 1), style=_history_text_c)
_record_output_history_entry(lambda: self._render_resume_history_panel_lines(panel))
with _suspend_output_history():
self._console_print(panel)