8523402db00a9f477dd2ee28cfaf974aec3ed555
416 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d532d83eca |
refactor(honcho): trim duplicated tests and long docstrings
Parametrize the sessionStart, injection-log, dashboard user_id, unresolved-peer and deferred-save tests that differed only in their inputs. Share the blocking remote in the concurrent flush tests. Fold _as_flag onto a word table. Cut the added docstrings to the what and the one non-obvious why. |
||
|
|
beab8b6f27 |
fix(honcho): keep observation flags across a flush rebuild, never orphan an evicted session, namespace dashboard logins
`_flush_session` discarded the observation flags when it rebuilt an evicted SDK session, and the cached path returned none, so recall fell back to the config snapshot. Both paths now return and store the flags. A deferred `save()` on a session the cap evicted puts it back in the cache, or flushes it inline when a newer object owns the key. `save()` and `stop_async_writer()` share the writer lock, and the writer drains its queue after the join, so a put that raced shutdown is written. The trim after a flush runs under the cache lock. The shutdown join takes the remaining budget instead of a fixed ten seconds. The injection audit file is created owner-only, and `logging: "false"` reads as off. The desktop passes `<provider>:<user id>` so a basic-auth alice and an OIDC alice are two peers. When a gateway platform supplies no user id, the peer notice and tool error no longer recommend peerName, which would merge every user of that gateway onto one peer. README documents `injection.sessionStart`, `logging`, and what a dashboard login does to peer resolution. |
||
|
|
c5973cd540 |
feat(tui-gateway): build the agent with the authenticated dashboard user as user_id
the dashboard login already stamps {user_id, provider} on the websocket
at upgrade time and tui_gateway keeps it as WSTransport.auth_identity,
but _make_agent never read it. every dashboard and desktop session was
built with user_id=None, so memory providers saw no runtime user and fell
back to the configured peer, mixing all logins together (#89794).
_make_agent now reads the session transport's auth_identity and passes
its user_id to AIAgent, the kwarg gateway platforms already use. the
legacy ?token= path, stdio, and the server-internal credential the PTY
child connects with carry no human and pass None, so they keep resolving
to the configured peer. the change is provider-neutral: honcho and any
other memory provider receive the id through the existing initialize()
kwargs.
|
||
|
|
11576390fe |
refactor(model): one persist writer for /model across CLI, gateway, TUI, dashboard; ACP + dashboard validate through switch_model
One `/model --global` produced four config.yaml shapes. CLI wrote
default/provider/base_url/api_mode and cleared the context pin on a route
change; the gateway rewrote the whole `model:` block (whole-file save_config)
and only set api_mode for `custom`; the TUI wrote three keys and never
touched api_mode, so a switch off an Anthropic-wire endpoint left a stale
`api_mode: anthropic_messages` in config; the dashboard main slot had its own
switched-provider logic, wrote `base_url: ""` and always dropped
context_length. ACP `session/set_model` and `POST /api/model/set` accepted
any model string (parse_model_input + detect_provider_for_model) so a model
no catalog knows, or a provider with no credentials, was handed to the
session / persisted and only failed at inference time.
Canonical: `hermes_cli.model_switch.model_selection_config_updates` (the
shape) + `persist_model_selection(result, config_path=None)` (targeted
per-key `atomic_roundtrip_yaml_update` writes, so sibling
`model_slots`/`model_fallback` keys survive; explicit path for the
multiplexed gateway's profile config) + `apply_model_selection` (same shape
applied to an in-memory `model:` dict for callers that save a whole
document). `atomic_roundtrip_yaml_update(value=None)` now REMOVES the key
instead of writing `key: null`, so per-key and whole-document writers land
the same file. Shape = CLI/gateway semantics: default, provider, base_url
(cleared when the target has none), api_mode (cleared when unresolved),
context_length cleared only when `should_clear_context_pin` says the route
identity changed, inline api_key/api cleared for non-custom targets.
Sites -> canonical:
hermes_cli/cli_model_switch_mixin.py::_persist_global_switch -> deleted; _commit_model_switch calls persist_model_selection
hermes_cli/cli_model_switch_mixin.py::_clear_persisted_context_for_model_switch -> deleted (folded into the shape)
gateway/slash_commands_model.py::_persist_model_switch_to_config -> to_thread forwarder: persist_model_selection(result, ctx.config_path)
tui_gateway/model_switch.py::_persist_model_switch -> deleted; _apply_model_switch calls persist_model_selection
hermes_cli/web_server_config.py::_apply_main_model_assignment -> apply_model_selection(result) (+ explicit custom api_key)
hermes_cli/web_server_config.py::_validated_main_model_selection -> NEW: switch_model(--provider) gate; rejection -> HTTP 400
hermes_cli/web_routers/{models,profiles,config_env}.py main-slot paths -> through _validated_main_model_selection
acp_adapter/server.py::_resolve_model_selection -> deleted; _switch_model calls switch_model (provider:model -> --provider), rejection -> ValueError
Behavior changes: TUI --global now writes/clears model.api_mode and clears a
route-changed context pin; gateway --global no longer rewrites the whole
model block (sibling keys survive) and clears api_mode for every target;
dashboard main slot / profile-create model / custom-endpoint activate now
reject unknown/uncredentialed/unlisted models (HTTP 400) and persist the
resolved base_url/api_mode instead of `base_url: ""`; ACP rejects the same
(ValueError surfaced by the command/protocol handler). Gateway persist runs
on a worker thread against the routed profile's config_path (multiplex-safe).
Cleared keys are removed from config.yaml rather than left as `null`. ACP
still never persists.
Kept `_normalize_main_model_assignment`: switch_model rejects a vendor name
posing as a provider (`moonshotai` -> "Unknown provider"), so the
vendor->aggregator repair is not a duplicate; E2E verified both branches.
No config migration: readers already coalesce `base_url: ""` to absent
(`_config_base_url_for_provider`) and gate api_mode on provider match
(`_provider_supports_explicit_api_mode`), so no stale-shape reader bug.
Tests: tests/hermes_cli/test_model_persist_one_shape.py (four surfaces land
one block; same-route re-pick keeps the pin), tests/acp_adapter/
test_acp_dashboard_model_switch_validation.py (rejection + explicit
provider prefix). Replaces test_acp_set_model_explicit_provider.py and the
two TUI-only persist tests; tests that intercepted the old per-surface seams
(`cli.save_config_value`, `load_config_readonly`, `tui_gateway.server.
_persist_model_switch`) now intercept the canonical seam. Each fix
sabotage-verified red.
|
||
|
|
9b9026e62f |
fix(kanban): drop the TUI empty-selection change and refit tests to the opt-in contract
Salvage follow-up to Xipong's #107736. Kept the core: `kanban` is a configurable, default-off toolset whose check_fn answers the schema build's own selection (ContextVar) instead of the legacy top-level `toolsets` key, so `platform_toolsets.<platform>: [.., kanban]` — what `hermes tools enable kanban --platform X` writes — actually reaches the gateway agent's tool schema. Dropped the `tui_gateway/server.py` change: turning an explicitly empty CLI selection from "all" into "nothing" is a separate behaviour flip already tracked by #107452, not part of this bug. The two TUI loader tests that asserted `kanban` is auto-recovered onto a saved `[memory]` list now assert the opposite: a configurable opt-in is never recovered. |
||
|
|
fc71fb63e5 |
test(multiplex): invariants for the residue fixes; retire allowlist tests
- tests/gateway/test_multiplex_residue_parity.py: served profile reads its own sessions.*; per-turn bridge skips secondary scope; resolve_proxy_url reads the routed scope (and does not borrow the default's on a miss); a stale served turn never recreates an archived profile; MCP discovery runs per profile home. - test_config.py: v43 migration removes multiplex_profile_allowlist. - Allowlist tests deleted (feature removed) or rewritten to "served set = all live profiles"; the unserved cases now use a tombstone / missing dir. - MCP discovery fixtures converted to the per-home set/dict slot. |
||
|
|
0dcadf6f41 |
revert: remove Collective Wisdom V1 (#94266)
Reverts the in-tree org skill-marketplace: hermes_wisdom package, three model tools, CLI/gateway/desktop/dashboard/Telegram/Slack surfaces. Later non-Wisdom work on shared files (guest onboarding i18n, dashboard startup schema, Slack adapter, tui_gateway) is kept; Wisdom-only call sites and config were stripped from those files. |
||
|
|
a6ee31f55a |
feat(wisdom): add Hermes Collective Wisdom Agent V1 (#94266)
* feat(wisdom): add trusted publish and install foundation
* feat(wisdom): add private contribution loop
* feat(wisdom): add managed consumption workflows
* fix(wisdom): close cross-repository safety gaps
* fix(wisdom): align local package and lifecycle policy
* fix(wisdom): require explicit profile setup
* docs(wisdom): repin reconciled gateway head
* fix(wisdom): fence content downloads and approval receipts
* docs(wisdom): record generation-fenced downloads
* docs(wisdom): record unified delivery PR
* fix(ci): stop passing invalid classifier inputs
* docs(wisdom): remove internal requirements ledger
* feat(wisdom): localize dashboard and desktop copy
* feat(wisdom): complete local contribution and consumption UX
* style(wisdom): satisfy desktop lint
* chore(wisdom): refresh requirements pin
* test(dashboard): allow formatted profile copy
* test(wisdom): stabilize desktop interaction coverage
* fix(wisdom): surface dashboard action failures
* fix(wisdom): add repeatable Portal demo login
* feat(wisdom): add actionable skill notifications
* feat(wisdom): add notification install and update actions
* fix(wisdom): make Telegram skill alerts actionable
* fix(wisdom): always refresh demo Agent login
* feat(wisdom): embed Telegram notification actions
* fix(wisdom): preserve Telegram notifications after actions
* fix(wisdom): keep Telegram notification cards readable
* feat(wisdom): add Telegram candidate approval flow
* feat(wisdom): explain Telegram qualification reasons
* fix(wisdom): reconcile cross-surface candidate actions
* feat(telegram): add Collective Wisdom management command
* chore(wisdom): refresh Gateway contract pin
* chore(wisdom): advance Gateway contract pin
* feat(wisdom): align command UX across clients
* feat(slack): add Collective Wisdom management parity
* feat(wisdom): add security and professionalism reviews
* feat(wisdom): add first-time qualification guidance
* feat(wisdom): simplify qualification sharing choices
* feat(skills): add optional editorial metadata
* feat(wisdom): enrich legacy skill presentation
* fix(wisdom): harden review and update boundaries
* fix(wisdom): emit canonical review timestamps
* fix(wisdom): align with merged gateway and main
* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)
- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
7-day evidence builder that excludes bundled/hub/managed skills and
dismissed/handled/recently-suggested content hashes, strict pydantic
schemas for agent output with repair-or-reject, fixed copy templates
(Share / Teammate / Published / Update / Mute), idempotent retried
delivery ledger with stale-action resolution, weekly review job,
resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.
* wisdom: agent-led renderers and button action dispatcher
- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
packaging flow, Install/Update -> plan command. Never publishes/installs.
* wisdom: CLI verbs, agent_led config default, conversational catalog skill
- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
verbs, share/install flows and fixed notification templates.
* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons
- gateway housekeeping tick calls maybe_run_weekly_review with a home
channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
duration keyboard, send_wisdom_agent_recommendation rich card + fallback.
* fix(wisdom): integrate local mediation and harden model and setup boundaries
* fix(wisdom): honor authoritative recommendation policy and defer on failure
* fix(wisdom): synchronize opaque suppression and recheck delivery preferences
* feat(wisdom): route weekly selection through the session-owned assessment queue
* fix(wisdom): prepare and submit the reviewed generated share package
* feat(wisdom): separate native Share preparation from publication consent
* feat(wisdom): sync native mute choices through a leased preference outbox
* feat(wisdom): bind native mute controls to durable preference choices
* feat(wisdom): add scoped desktop and dashboard notification settings
* fix(wisdom): revalidate feed recommendations before assessment and delivery
* fix(wisdom): persist validated delivery receipts before completing notices
* feat(wisdom): add private notification claim and receipt client
* Persist Wisdom send reservations and recover delivery acknowledgements
* Route legacy Wisdom controls through current native review
* Add typed private Wisdom operation outcome client
* fix(wisdom): make agent-led advice usable in the local demo
* fix(wisdom): keep requested consent outside proactive limits
* fix(wisdom): distinguish unavailable assessments and preserve digest text
* fix(wisdom): assess ongoing usefulness beyond the current task
* fix(wisdom): restore immediate qualification sharing controls
* fix(wisdom): separate qualification review from installation advice
* fix(wisdom): collapse review checklists and simplify sharing copy
* fix(wisdom): show compact sharing progress and publication receipts
* fix(wisdom): require credential prefixes rather than matching skill names
* fix(wisdom): finish package checks before presenting sharing consent
* fix(wisdom): scan local skills before qualification cards
* fix(wisdom): update moderation results on existing sharing cards
* fix(wisdom): keep sharing review accessible from receipt cards
* fix(wisdom): align mediated review cards and collapsible checks
* fix(wisdom): clarify clean security summary wording
* fix(wisdom): normalize consent plans and add explicit recheck
* fix(wisdom): keep install and update receipts concise
* fix(wisdom): collapse assessments and deduplicate operation cards
* fix(wisdom): restore private Portal review from native cards
* fix(wisdom): sync Portal publication to original consent card
* fix(wisdom): show local skill version on sharing cards
* fix(wisdom): skip agent recommendations for self-published versions
* fix(wisdom): simplify candidate notices and local-edit recovery copy
* feat(wisdom): submit locally reviewed packages with one confirmation
* feat(wisdom): expose safe receipt and outcome sync recovery
* wisdom: onboarding notice says detect and share, names the user's own skill
Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark
Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.
* wisdom: one opener, no approval line, ask to share after the skill is shown
Product owner review of the candidate card.
- The Hermes written card now opens with the same sentence as the fixed card
("Your organisation has enabled Collective Wisdom, a feature designed to
automatically detect and share useful skills across all team members.")
instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
It is now the last line, after the skill name, description, why suggested
and the checks, and reads "Would you like to share it?" (matching the
agent led template wording).
Tests updated for the new order; proposalNotice removed from all desktop locales.
* wisdom: American spelling, organization
Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.
* wisdom: candidate card copy round 4 (owner review)
Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:
1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
card (Telegram rich card and plain fallback, legacy agent-led share
template).
3. The skill name and description are labelled: "Skill name: <name>" and
"What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
inappropriate content found)" with no per-check bullets and no "Pass";
a failed review reads "Needs a look before sharing at work (possible
inappropriate content)" and lists only the checks that flagged
something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
editorial_name, a simple one_line_description and a compelling
why_coworkers_benefit under 300 characters; "Be concise and
convincing." becomes "Be concise and compelling: the goal is that the
user wants to share it."
Tests updated for the new strings; review_text() gains direct coverage.
* wisdom: re-apply owner copy after rebase
- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice
* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors
Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.
* fix(wisdom): reconcile optional SDK tests and frontend lint
* fix(wisdom): default to agent-written notification summaries
* fix(wisdom): restore deferred install review and browse controls
* feat(wisdom): inspect installed setup with exact package provenance
* feat(wisdom): run native-approved installed setup steps with durable evidence
* fix(wisdom): recover interrupted setup with explicit native consent
* feat(wisdom): hand native installs into guided setup review
* fix(wisdom): continue requested setup with fixed notification copy
* fix(wisdom): preserve setup while waiting for a session model
* fix(wisdom): expose canonical setup review controls on desktop
* fix(wisdom): resume setup after recorded automatic updates
* fix(wisdom): make missing setup prerequisites recheckable
* chore(wisdom): align Agent with verified Gateway contract
* fix(wisdom): stop guessing team slugs in portal links
* fix(wisdom): retire pending advice on account sign-out
* fix(wisdom): cancel advice after terminal account revocation
* fix(wisdom): fence feed responses across account sign-out
* fix(wisdom): checkpoint signed-out feed before reactivation
* fix(wisdom): link proactive advice to scoped notification settings
* fix(wisdom): coalesce queued publication recommendations by version
* fix(wisdom): keep package review navigation local and deferable
* fix(wisdom): reflect installed state in discovery controls
* fix(wisdom): show exact checks before command confirmation
* chore(wisdom): pin bounded analytics privacy contract
* chore(wisdom): pin retired legacy notification contract
* feat(wisdom): review publisher usage with exact sharing copy
* fix(wisdom): align discovery and review check summaries
* fix(wisdom): show expired consent before confirmation
* fix(wisdom): require fresh review for legacy install controls
* fix(wisdom): preserve review expiry across check toggles
* fix(wisdom): retain update policy in native install reviews
* fix(wisdom): surface failed native card edits
* fix(wisdom): persist local command approval reviews
* fix(wisdom): use saved approvals for messaging commands
* test(wisdom): provide scan result in setup handoff fixture
* test(wisdom): exercise Telegram approvals with saved review state
* fix(wisdom): retain suppression policy for offline deferral
* fix(wisdom): reconsider candidates after deferred suppression expires
* fix(wisdom): bind review checks and report verified readiness separately
* fix(wisdom): persist accepted publication intent and recover exact outcomes
* fix(sync): pin UTF-8 tree ordering across writers
* chore(wisdom): pin organisation-scoped Gateway authorization
* fix(wisdom): restrict consent delivery to user-facing sessions
* chore(wisdom): refresh reviewed Gateway contract pin
* fix(wisdom): preserve kept tools in Blank Slate exclusions
* test(auth): reset anonymous fixture with a profile-scoped cache
* fix(wisdom): gate local surfaces and work on current profile entitlement
* fix(wisdom): invalidate quiet tool cache on entitlement changes
* test(wisdom): authorize local consent gateway fixtures
* fix(wisdom): keep entitlement decoding free of native crypto imports
* test(wisdom): provide local entitlement to demo CLI subprocess
* ci: leave upstream workflow unchanged in Wisdom PR
* fix(wisdom): ship package and contracts in Nix wheels
---------
Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
|
||
|
|
4bdd64b334 |
The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an anonymous account exactly one model on its own host, refuses everything else with a structured 429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and tells a signed-in account that still asks for `nous/welcome` what to switch to in an `x-nous-model-switch` header. Four client-side gaps against that contract: - Auxiliary calls were refused on every session. The auxiliary client asked the welcome host for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free` before each fallback. On the welcome host it now uses `nous/welcome` (its backing model covers auxiliary work) and skips Nous for vision, which the welcome model does not take. - The structured 429 body was never read. The classifier now parses `reason` / `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited` are rate limits that honour `retry_after` and never rotate the free tier's only credential. The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and fall back instead of retrying or re-exchanging. The terminal paths say what happened and name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal). - The `x-nous-model-switch` header was ignored. The chat-completions transport records it beside the rate-limit and credits headers; the next call moves the session, and the config default when it still names `nous/welcome`, to the backing model the gateway named. - A guest fell back to the paid host. With `inference_base_url` absent from the exchange or outside the host allowlist, routing defaulted to inference-api, where every request is a 400. A guest now defaults to the welcome literal at the exchange, in the shared store's shape, and in effective routing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8) * feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime rung now sets it up there instead of failing "not logged in", so the guided chat no longer races the root profile's first-run mint. nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever nothing else is configured; "on-request" mints only when the free tier is asked for by name (nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers — the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the connector token path — still adopt what the shared store holds, so every profile follows the one identity the guided setup created, but never create one on their own. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c) (cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165) * feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity (nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is configured. "explicit" means Hermes never creates one on its own: the only creator is the new provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and before the guided chat exists — so the identity lands in the root store every profile reads through and is there before any session asks for nous/welcome. That closes the race against the backend's own setup, and makes "only when the setup-bot flow is used" literally true. The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome (the hermes model row, --provider nous), which treated a model name as intent and was broader than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to sign in from. Implicit callers still adopt an identity the shared store holds, and a retired credential is replaced (a continuation, not a creation). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0) * fix(auth): remove the nous.guest_setup knob; the free tier is created on first use `nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`, unknown values read as `auto`, and under `explicit` a CLI-only install could never get an identity, which contradicts the first-run contract (first command mints, then chats). The mint race the knob accompanied is already benign: every caller takes the profile lock then the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which stays. `nous.guest` remains the only free-tier policy. Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through `ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config entries. The three policy tests that hold regardless of the knob are kept under `TestExplicitProvision`; the two that only tested the knob are deleted. (cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261) * fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch` moves the live session to that model and moves `config.yaml`'s default off the alias in the same step. The messaging gateway's fallback-eviction check compares the agent's model with the config default and evicts on any mismatch that is not a /model override, so when the config write did not land (unreadable config, lock) the cached agent was evicted once per turn, and prompt caching with it. `apply_model_switch` now stamps the alias it moved the session off on the agent, and `_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as deliberate, beside the existing /model override case. The check takes the agent and the config model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them. (cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954) * fix(auth): the free tier outranks implicit host credentials in provider resolution On a fresh install with a leftover ~/.aws profile, resolve_provider("auto") reached the Bedrock rung before the free-tier rung, so the first turn ran on Bedrock and failed 403 while the free tier was still being minted in the background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s, three retries, no answer; the next process then switched to nous/welcome. The free-tier rung now sits directly above the Bedrock chain: when nous.guest is on, an existing free-tier identity answers, else a blocking mint runs, and only then does the boto chain get a say. Everything above is unchanged and still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter pool, a logged-in active_provider. nous.guest: false skips the rung, and a failed mint still falls through to Bedrock and the no-provider guidance. Tests: six precedence cases (identity present, fresh mint, free tier off, env key still wins, sign-in still wins, failed mint falls through). The opt-out test now neutralizes the AWS chain like the precedence tests do; on a machine with ~/.aws it was failing for the same reason as the bug. Live after the fix, same Mac, AWS credentials visible, isolated shared store: identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in 11 s. (cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a) * fix(auth): review follow-ups for the free-tier rung (NS-829) - tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches the free tier off; its contract is the boto chain, and the free tier now sits above it. - gateway/run_notifications.py: the free-tier startup line reads auth.json before consulting the resolver, so a gateway boot on a machine with AWS credentials never mints or refreshes over the network. - hermes_cli/anon_auth.py: module docstring says where the free tier sits in the ladder instead of "the ladder is untouched". - tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of six (parametrized ladder cases; a failed mint that returns None or raises falls through to Bedrock). scripts/run_tests.sh on the five affected files: 147 passed, 0 failed. (cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11) * feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone The free tier is pre-GA. Until GA it must not exist for anyone who did not ask for it: no identity minted, no portal traffic, no free-tier copy on any surface. One environment variable now decides that, and one function reads it. `guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly "1"; only then does `nous.guest` (the user's off switch) get consulted. Every free-tier site already funnels through `guest_enabled()`, so the gate closes minting, routing, connector entitlement, status lines and the picker row in one place. With the variable unset, `resolve_provider("auto")` on a fresh install raises `no_provider_configured` exactly as upstream does. `HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted identities as a side effect of provider resolution, and `_has_any_provider_ configured` read them ahead of every other check, making the CLI a second reader of a flag that must have exactly one. `_forced_new_done` and the `force` parameter of `_reconcile_and_provision` go with them. Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847). Not a user preference: the variable is never written to config.yaml or .env and never shown in setup. It is deleted at GA together with its comment in anon_auth.py. This is a deliberate, temporary exception to the "no new HERMES_* env vars for non-secret config" rule. Tests: fixtures set the gate instead of deleting the old lever; one new invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "", "0", "true" and "new" all leave the tier off with zero portal calls, red on the previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the feature. * feat(auth): the free-tier identity is created in one place, at boot; every other site is a read Before this commit eight sites could create a Nous free-tier identity as a side effect of something else: resolving a provider, the CLI's first-run check, the CLI's session setup (in the background beside an own key), a connector bearer read, the desktop polling `free_tier.status`, the sign-in precondition, the desktop's `free_tier.provision`, and the dead-credential re-mint. A poll could mint. Provider resolution could hit the network. Two of them raced each other on a fresh install. Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator. `hermes serve` runs it on a daemon thread from `_lifespan` beside the other background boots; `cmd_chat` runs it synchronously before the first-run guard. It inventories credentials first (`resolve_provider("auto", skip_free_tier=True)`: what would carry inference if the free tier did not exist), creates the identity only when `guest_enabled()`, resolves inference, records a `SetupRecord` in process memory and broadcasts ONE `setup.ready` event. It runs on every boot; only the mint is gated. `ensure_portal_identity` now requires `explicit=True` and raises otherwise. Its callers are the bootstrap, the desktop's `free_tier.provision` (the explicit retry when the boot could not create the identity) and the two dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`, `managed_tool_gateway._replace_dead_guest_token`). The background thread path and `provision_free_tier` are deleted with their last callers. Reads that used to mint and now only read: `auth.py::resolve_provider` rung 7 (an existing identity still outranks the Bedrock chain, NS-829 ordering kept), `main.py::_has_any_provider_configured`, `cli_agent_setup_mixin._ensure_runtime_credentials`, `managed_tool_gateway.read_nous_access_token` (no identity -> None), `anon_sign_in.run_sign_in` (no identity -> Unavailable), `methods_free_tier` `free_tier.status`. `setup.status` answers from the record for the launch profile, blocking up to 8 s while the bootstrap is in flight so a client's first poll lands after the identity exists rather than racing it; a named profile, or a process that never ran the bootstrap, keeps today's live probe. The record's fields ride along additively (`ready`, `free_tier`, `other_providers`, `inference_provider`). Identity and inference are decoupled (NS-845 Q1.3): the mint sets `active_provider="nous"` only when the inventory found nothing else usable (`_mint_locked(carries_inference=)`); an adopted account always does. A token refresh no longer re-elects the provider it refreshed (`_save_provider_state_to_source` writes credentials, not the user's choice) — that write was how an own-key install ended up on the free tier after the first connector call. Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check), bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint), 62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and `provision_free_tier`), and a04b05260c (blocking mint in the resolver). Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847. Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps inference; reads never reach the portal; a refused mint is memoised), `free_tier.status` fails loudly if it ever calls the creator, the resolver stub fails loudly if resolution ever mints, `setup.status` reads the record, `skip_free_tier` proves the inventory question. The three sign-in tests for the deleted pre-mint collapse into one (`no identity -> Unavailable, zero portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20. * fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup" A free-tier identity carries $0 by design, so the portal seed reports `paid_access=False` for it. `is_free_tier_model` did not know the welcome host, read that as a depleted account, and every free-tier turn ended with the credits-depleted notice telling the user to top up an account they do not have. Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host (`anon_auth.route_is_welcome_host`) is the free tier. The host is the evidence, not the model name: the paid inference host can serve `nous/welcome` to a named account and that account's depletion is real, so `("nous/welcome", <inference host>)` stays False. Local data only, like the three rules above it. Restores the two contracts dropped by hermes-magic 674e11d1eaa (the prototype line ran without unit tests): the welcome host is free without any pricing evidence; the model name alone is not. The first is red without this fix. * fix(copy): free-tier text stops promising a connector transfer and never names the config key Sign-in copy on every surface said "Sign in to keep your connectors" and ended with "Your connectors are kept." The transfer registry that would make that true is empty (NS-821): nothing carries over today. The copy now says what signing in does give ("unlock more models and tools") and the completion line names the account, not a transfer. The docs page loses the "connectors carry over" paragraph for the same reason. The picker's off-state line exposed `nous.guest: false` and the word "guest"; user copy names the free tier only (R-USR-1). The docs page gains the pre-rollout note: until GA nothing on it happens without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and "replaced on next use" sentences now describe the boot bootstrap. zh is a strict locale: the `freeTier` block was English placeholder text copied from `en`; it is now Chinese. `connectorsKept` is renamed `completedBody` since it no longer talks about connectors. * feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1" as on. Until now nothing in the desktop set it, so a packaged app could never turn the free tier on, and a backend spawned by the app could disagree with the app about whether the tier was live. `electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled` is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has `--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch into a module constant. `desktopBackendSpawnEnv` wraps every backend env as the outermost call and writes the flag LAST, as "1" or an explicit "0", so no earlier spread (`process.env`, `backend.env`) can resurrect a stray value from the parent shell. Stamped onto all three spawn sites: the primary `serve` spawn, the pooled per-profile spawn, and the remote SSH `exec env ...` command (which gains ` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and the backend probes are not backend spawns and do not get it: a `hermes --tui` typed in the pane must not mint. The renderer learns the same fact read-only through the existing `hermes:launch-flags` sync IPC (`guestOnboarding`) and preload (`window.hermesDesktop.guestOnboardingEnabled`). Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding` maps to it in main). Two invariant tests on the pure helpers: only "1" or the argv flag enables; the spawn env carries "1"/"0" as the last word and preserves every other key. * feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll The backend's boot bootstrap now announces `setup.ready` once, after it has created (or refused) the free-tier identity and resolved the inference route. The renderer used to discover both by polling `setup.status`, `setup.runtime_check` and `free_tier.status` every 60 s from `useStatusSnapshot`; a fresh install's chip, notice strip and onboarding overlay could sit stale for up to a minute after boot, and three RPCs a minute per window kept asking a question whose answer changes only at boundaries the backend already announces. `handleLifecycleEvent` routes `setup.ready` (active source only, like `skin.changed`) to `notifySetupReady()`, a one-shot tick atom in `live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens to it and runs one readiness round at once (`setup.status` + `setup.runtime_check` + `free_tier.status`). The readiness legs also run once on open and on return from another app, as today. The 60 s tick keeps only `getStatus()`. `SetupStatusSnapshot` types the record's additive fields (`ready`, `free_tier`, `other_providers`, `inference_provider`); readiness semantics are unchanged and still key on `provider_configured` + `runtime_check`. Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one refresh from the active source and none from another; the snapshot hook's contract is three legs on open, one leg on the tick. * fix(cli): the banner names the free tier's model instead of "no model configured" The welcome banner prints before credentials resolve, so on a fresh install `model` is empty and the banner said, in red, "no model configured — run /model or hermes setup". Under the free tier that is false: the route is already known from local state (identity on disk, tier on), and the first message will run on `nous/welcome`. `_banner_left_lines` now asks the route the same question when `model` is empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the banner's 'no model configured' line reads the resolved route"). Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`; gate off -> the red line, zero portal calls. * fix(aux): vision on the free tier uses nous/welcome too The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the welcome host would have sent every image step past the free tier for no reason, so the auxiliary client pins the route's one model for every lane. A backing model that takes no images answers with the upstream's own error, which the ladder handles as it always has. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415) * fix(gateway): hermes gateway run is a boot owner of the free tier too Rung 5 made every demand-time free-tier site a read: resolve_provider, the connector token, the /login precondition. That is only correct if every process that can reach those sites ran the bootstrap first. The CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone messaging gateway did not. A fresh HERMES_HOME with the gate on and `hermes gateway run` reached provider resolution with no identity to consume, and /login returned Unavailable. Reported by @andrexibiza on #107697 (P1). GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an executor thread right after startup recovery and BEFORE any adapter connects, so a fast first DM cannot arrive with nothing to resolve. It is its own step, not part of the turn-machinery warm-up: the warm-up is an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0); the bootstrap is correctness and must always run. With the gate unset it is a local inventory and no network. Live, real GatewayRunner.start against a fake portal in a fresh home: gate on -> 1 create, identity persisted, resolve_runtime_provider=nous, /login precondition sees the identity gate off -> 0 portal calls, no identity, no_provider_configured Before the fix the gate-on row was identical to the gate-off row. Test: the bootstrap seam runs before _start_prefilter_platforms and delegates to the one creator. Red on 5554eb6993 (no seam), green here. --------- Co-authored-by: Robin Fernandes <robin@soal.org> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
c22a8d8e3f |
Seeded sessions survive a gateway restart and store their seed once (tui_gateway) (#107549)
* fix(tui-gateway): a seeded session is durable at create, and its seed is written once session.create accepts opening messages. Three defects sat in that path: - A seeded session without a parent was never persisted at create, so a restart before the first prompt lost it and session.resume answered 4007. Only branch children (#93959) were persisted up front. The same rationale applies to any seeded create: seeded content is intent, not an abandoned draft. Parentless seeds now persist their row, transcript and client title at create; empty drafts stay lazy. - _coerce_seed_history dropped display_kind, so a seeded row tagged "hidden" (model-facing scaffolding) rendered as a user bubble. The coercion keeps "hidden" and only "hidden"; every other kind is stamped by the gateway at turn time and is not accepted from the wire. - A branch child's seed was written twice: _seed_branch_row copied it at create but never marked it persisted, so the first prompt's _persist_branch_seed appended the copy again. The create path now sets _branch_seed_persisted, and the gate is a create-time `seeded` stamp instead of parent_session_id, so a resumed session (whose history comes from the DB) can never re-append its transcript. Two invariant tests, both red on main: a parentless seed survives a gateway restart with the hidden row kept out of the wire transcript and not re-written by the first-submit path; a branch child's seed is stored exactly once. The reasoning-fields fixture stamps `seeded`, the flag session.create sets. * fix(tui-gateway): a hidden seed row stays out of the list preview and the create count Live-testing the seeded create on every surface showed two places where the newly durable hidden row (display_kind="hidden") still surfaced: - session.list built a session's preview from its first user row with no display_kind filter, so a hidden opening row (model-facing scaffolding the gateway never paints) became the sidebar preview. The preview predicate now skips hidden rows, in every listing query that shares it. - session.create reported message_count as the raw seed length while its messages array already filtered the hidden row (2 vs 1). It now counts what is on the wire, the same rule session.resume applies. Both are covered by the existing seeded-create test: the create count equals the wire transcript, and the preview of a session whose first user row is hidden is its first visible user row. * fix(tui-gateway): a live unpersisted resume counts the wire transcript session.resume on a live session that has no row yet reported message_count as the raw history length while its messages array was already filtered, the same mismatch the previous commit fixed on session.create. Count the wire, as the cold, deferred and reuse-live resume paths already do. * chore: retrigger CI (zero-job dispatch failure, auto-heal) |
||
|
|
75b1b43ec1 |
fix(tui-gateway): reject sessionless config.set model with 4001
Sessionless model switches (including legacy --global) could persist profile defaults before session.create; fail closed unless --once. Co-authored-by: Cursor <cursoragent@cursor.com> (cherry picked from commit 1c4a29757bbde1304404f26b45b6becb841c069f) |
||
|
|
06dc51d62d |
fix: verification evidence ledger is inert while verify_on_stop is off
The ledger in verification_evidence.db exists only to feed the verify-on-stop guard, but the recorder kept running on every foreground terminal command and every file edit after #53552 turned the guard off by default. Users who never opted in still accumulated a multi-MB database (7 MB / 4.6k rows on one install). Every ledger entry point (record_terminal_result, record_verify_run, mark_workspace_edited, verification_status) now checks verify_on_stop_enabled() first and returns without opening or creating the database when the guard is off. verification_status reports {"status": "disabled"} in that case; no client consumes the verification.status RPC yet, so nothing downstream changes. Existing ledger tests pin HERMES_VERIFY_ON_STOP=1 since they exercise the ledger itself; the new test proves the off path never creates the file (red on base). |
||
|
|
d095f8fb15 | fix(tui): bound subscriber delivery without blocking shared turns | ||
|
|
de25545dce |
fix(tui_gateway): fan session events out instead of rebinding the transport slot
A session held exactly one transport, and prompt.submit, session.resume, session.activate, and the queued-prompt drain all rebound that slot. A second client therefore took the stream away from the first: the earlier client stopped receiving the turn it was already rendering, and either client disconnecting parked the whole session on the drop sentinel. FanoutTransport goes in the same slot and satisfies the same Transport protocol, so write_json and every other reader of the slot are unchanged. It delivers each frame to a snapshot of its peers, concurrently when more than one peer is attached and the caller is not on an event loop, and prunes any peer that returns False or raises. A dead client is dropped; a slow one costs the emitter at most one write timeout per frame rather than one per peer. Request/response RPCs are unaffected: they still answer on the request's context-bound transport, so a client only ever sees replies to its own calls. The rebind sites become attach sites through _attach_session_transport, whose ladder keeps the single-client shape identical. The same object already in the slot is a no-op; an empty, stdio, or parked slot is taken outright; only the arrival of a second live client wraps both. The queued-prompt drain is included because it pinned the drained turn to the queuer and silenced everyone else. A non-peer newcomer such as stdio or the drop sentinel never displaces a live client, so an activate dispatched without a bound websocket cannot silence the socket that owns the session. Disconnect detaches first. A session that retains another client keeps streaming and is neither parked nor reaped, and only the clientless ones follow the existing close_on_disconnect and park-sentinel path, so a single-client disconnect, the orphan reaper, and its grace window behave as before. _ws_session_is_orphaned is unchanged: it still asks whether the drop sentinel is in the slot, and a fan-out is never the sentinel, so a session that still has a peer is never reported as orphaned. Attach performs no entitlement check: any authenticated peer may mirror any session. The fan-out architecture follows the approach in #40822 by @OmarB97. What the slot's later history forces. _close_sessions_for_transport drops the #83716 rebind-to-the-most-recent-surviving-viewer, which fan-out membership subsumes — a pop-out window is a peer, so a session that still shows in one is never returned as clientless — and keeps the #77129 revalidation before parking, now expressed as a liveness check under _session_transport_lock so it is race-free against attach and detach. _transport_is_live_peer defers its last answer to _transport_is_dead: a socket that already latched _closed is a departed client, and admitting it would keep a session out of both the park and the reap. _transport_is_dead also learns the fan-out: a FanoutTransport with no live peer is dead, so a session whose peers were all pruned by failed writes cannot outlive the TTL and LRU reapers. The upstream test that pinned the #83716 rebind, test_close_transport_rebinds_session_to_remaining_viewer, is re-expressed in fan-out terms: both windows attached, the pop-out closes, the session stays with the main window unparked and still receiving frames. |
||
|
|
7f80efc294 |
fix(tui): post-turn completion drain resolves ownership once; barrier test uses a watch
Rebasing onto main surfaced three reds in tests/test_tui_gateway_server.py. Two were a real double-check: _run_post_turn_followups already drains the completion queue through _session_owns_notification_event, then handed the events to _notif_handle_ready, which re-ran the belongs-elsewhere and requires-owner gates on every event — a second compression-lineage DB lookup per notification for events the caller had just proven ours. _notif_handle_ready/_notif_handle_event take an explicit owned=True from the post-turn path and skip only the ownership gates (consumed/dedup/turn claiming still run). The poller path is unchanged (owned defaults False). The third, test_run_prompt_submit_requeues_all_unstarted_notifications_with_real_threading, asserted that three consecutive completions produce one in-flight turn plus two requeued events; under this PR's contract (#104671) consecutive completions share one turn, so the first event is now a watch_match — the documented turn barrier — and the test keeps proving that the two completions behind an in-flight notification turn are never lost. A/B: revert the tui_gateway hunks and the two lineage tests fail with ownership_checks recorded twice; restore and all 11 test_run_prompt_submit tests pass. |
||
|
|
ab1a6d016a |
test(display): encode provenance contract in tilde assertions
Two tests still encoded the pre-PR behaviour that the PR removes:
- tests/test_tui_gateway_server.py: the mirrored fixture carries no
`context_estimated` flag, so under the PR's rule it is provider-reported
usage and must render without `~`. The old expectation asserted the
unconditional tilde main used to emit. Assert the flag-less case is
unmarked and, in the same test, flip `context_estimated` both ways to
pin that only the estimate carries `~` in the count and the percent.
- ui-tui appChromeStatusRule.test.tsx: `text.includes('~')` matched the
`~/repo` cwd label, so the "not estimated" arm was always true. Extract
the rendered context token and assert the tilde on that token only.
|
||
|
|
026e3e84ea | fix: reject unavailable desktop profile and session targets | ||
|
|
8513c5984c |
[verified] fix(gateway): fence orphan timers and reconnect interrupt claims
Re-scope #98106 onto the activity-based orphan policy from #100504. Keep timer ownership across callbacks and continuations, serialize reconnect paths against interrupt claims, and avoid recursive eager-resume locking. Leave cleanup polling armed when concurrent cold reuse is rejected. |
||
|
|
deb9f3c1bd | simplify(compat): hermes_cli/commands — drop the last PEP 562 lazy re-export hook (25 names), repoint 10 callers + 11 test files to commands_platforms/commands_completion | ||
|
|
310bed4302 | fix(compat-fallout): tui_gateway tests stub the defining modules (tools.browser_tool_lifecycle, tools.tts_tool_speaker) instead of the old tools.browser_tool/tts_tool facades | ||
|
|
7a33369e81 |
simplify(compat): interrupt — drop _ThreadAwareEventProxy/_interrupt_event legacy alias, repoint 2 test files
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt() directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker. |
||
|
|
c93ace77c2 | simplify(compat): config/runtime_provider/plugins/commands/secrets_cli/kanban — drop 96 re-exports (incl. PEP 562 facades) + 3 aliases (get_pre_tool_call_directive/_block_message, get_telegram_handler_factories), repoint 56 callers + 50 test files | ||
|
|
53db597201 |
simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by
|
||
|
|
0515997887 | simplify(compat): voice_mode/transcription_common — drop 9 re-exports + the _configured_stop_phrases seam, repoint 3 callers + 4 test files | ||
|
|
f6938b37f3 |
simplify(compat): terminal/file/environments — drop 42 re-exports/aliases, repoint 20 callers + 41 test files
tools/terminal_tool.py: drop 30 pure re-export names (lifecycle/config/backends/
sudo/guards/result/interrupt/utils/_DockerEnvironment/is_managed_tool_gateway_ready)
and the noqa-F401 comments on the 25 names the facade itself uses. Sibling modules
(terminal_tool_backends/_result/_sudo/_lifecycle, environments/base, process_registry)
that read removed names through the facade now import from the defining module.
tools/environments/base.py: drop 11 re-exports (base_output/base_session_env/
path_utils) and the BaseEnvironment.stop() compat alias (no in-tree caller; the
lifecycle hasattr(env, 'stop') fallback stays for third-party envs).
tools/environments/docker.py: drop 1 re-export + the re-export comment.
Callers/tests repointed to tools.terminal_tool_{lifecycle,backends,sudo,config,
guards,result}, tools.interrupt, tools.environments.{base_output,base_session_env,
path_utils}.
|
||
|
|
98c140bc4b | simplify(compat): code_execution_tool/environments.local — drop 36 re-exports, repoint 5 callers + 14 test files | ||
|
|
eeb7671e69 | simplify(compat): hermes_cli small facades — drop 7 re-exports/aliases (+relay_runtime alias module), repoint 12 callers/tests | ||
|
|
16370ae539 |
fix(tui_gateway): setup.status / setup.runtime_check answer for the requested profile
Port the Python half of PR #94147: both readiness RPCs accept an optional `profile` and bind that profile's HERMES_HOME + .env secret scope for the duration of the check via `_session_profile_runtime_scope` (ContextVars, so concurrent checks stay isolated). Unknown profile → ok=False with an explicit error instead of quietly reporting the launch profile's readiness. `_has_any_provider_configured(strict_profile_scope=True)` reads provider env only from the bound secret scope (never os.environ) and skips the host-wide fallbacks (gh auth, Claude Code credentials, api-key active_provider in auth.json) that describe the launch host, not the target profile. Unscoped callers are byte-identical to before. The desktop TS half of #94147 targets plugin.js, which was deleted on main; it needs a recut on create-dialog.tsx. Supersedes #94147 (python half) Co-authored-by: Zeus-Deus <100132710+Zeus-Deus@users.noreply.github.com> |
||
|
|
8076c78c87 | fix(desktop): scope handoff config to session profile | ||
|
|
000d6efd09 | test: widen _find_live_session_by_key monkeypatch for the optional profile_home arg | ||
|
|
c7e2e0b779 |
feat(fast): bounded /fast auto|cold windows behind one route-aware gate
Adds two bounded fast modes on top of the static /fast toggle, default OFF: - `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s) window; requests inside it carry the provider fast param, later tool-loop requests fall back to standard pricing. - `cold`: the same window, but only on the first turn of a session (no prior user/assistant/tool history). agent/fast_mode.py holds the whole policy: `begin_turn()` at the run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()` is consumed in the ONE place request_overrides feed the transports (build_api_kwargs), so the fast param is a per-request kwarg only. System prompt, tools and messages are untouched — the prompt cache is preserved. resolve_fast_mode_overrides() is now the single gate for static and bounded modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure, Bedrock and custom base_urls never receive service_tier/speed (#34308's route gating). Both existing callers (CLI turn route, gateway turn route) and the TUI config.set path pass the route. Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`, `/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop config.set; status shows the mode; web dashboard select lists the real values. Docs: configuration.md Fast Mode section with mode table + cost note, slash-commands, cli-config.yaml.example, locale strings for the two picker entries. Salvages #89991 (bounded fast modes) and #34308 (route gating). Fixes #64785, #74730. Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com> Co-authored-by: kbaicai <kbaicai@qq.com> |
||
|
|
aa80626764 |
fix(tui-gateway): adopt late compute-host compress acks instead of a false 120s timeout (#97948)
Manual /compress on a compute-host (turn_isolation) session blocked its RPC waiter for a hard-coded 120s, answered error 5019, and then DROPPED the host's late `control.ack`: HostSupervisor.control() popped the pending queue in `finally`, so `_handle_host_frame` had nothing to deliver to. The host kept compressing, succeeded minutes later, rotated the session — and the gateway session never mirrored the new session_key/history_version and the desktop never refreshed its transcript. - host_supervisor: `control(..., on_late_ack=)` leaves a one-shot handler registered when the waiter times out; control.ack/control.error/error frames for that request_id fire it (bounded: 30min TTL, cap 64). A host crash fails outstanding handlers with a synthetic control.error. - server: `_compute_host_compress_wait_seconds()` derives the wait from `compression.context_total_ceiling_seconds` (+30s slack, floor 120s, cap 630s) instead of the literal 120. `_adopt_late_compute_host_compress_ack` applies the metadata mirror and emits the same `session.info` a normal compress does plus the existing `status.update kind=compacted` edge; a late error goes out through the existing `error` event. - session.compress / slash.compress (methods_tools + _mirror_slash_side_effects): on waiter timeout answer `status: pending` (not 5019) and register the late-ack handler. - desktop: SESSION_COMPRESS_TIMEOUT_MS 120s -> 660s (above the gateway cap); `status: 'pending'` renders as an info notice, not `error:`; the `compacted` status edge rehydrates an idle active session's transcript (mid-turn compaction still defers to the turn settle path). Minimal extraction of the design in #99630 by @vsd2807 (design trace by @andrexibiza and @JoaoMarcos44 in the #97948 thread); no new DB tables, modules, or polling protocol. Refs #97948 Co-authored-by: VVV <vaibhavdahiya28@gmail.com> |
||
|
|
78eb4ebbfc |
fix(gateway): compensate half-written seeded branches + observable failures
Review follow-up on #93959: 1. Partial-failure window: if the row commits but the transcript copy or title write fails, the durable-but-empty child defeated the lazy first-prompt fallback (_ensure_session_db_row is INSERT OR IGNORE), so the renderer fail-latched on a transcript-less session again. The seed block now compensates: delete just this child so the lazy path can retry cleanly. Disk-full is exempt — deleting data on a full disk makes things worse. 2. Silent degradation: the best-effort catch now logs at WARNING with exc_info instead of DEBUG, so a regression in this user-facing path is observable without enabling debug logs. Tests: compensation deletes the half-written row and preserves pending_title; disk-full keeps the row and surfaces the WARNING. |
||
|
|
6556439765 |
fix(gateway): persist seeded branch children at session.create (#93959)
Desktop branch creation hung on an infinite spinner and lost the branch on restart. Root cause: the renderer branches via session.create with parent_session_id + a seeded transcript, but session.create defers the DB row to the first prompt (the draft-hygiene contract). The renderer's post-create resume then re-fetches the fresh child through REST and defer_history hydration — both read the DB. An unpersisted child 404s and hydrates empty, the client fail-latch (sessionShouldHaveTranscript + empty messages) refuses to bind a "transcript-less" session, and the user sees a spinner forever; on restart the rowless child vanishes and the optimistic "Draft: Branch N" entry disappears with it. A seeded branch is explicit user intent, not an abandoned draft. session.create now persists the child immediately when both parent_session_id AND seeded history are present: - Row created in the PARENT's profile-scoped state.db, stamped with _branched_from + parent_session_id (same shape as TUI /branch). - Seeded transcript copied via append_messages_batch so REST prefetch and defer_history hydration find it on the first read. - Title assigned from get_next_title_in_lineage(parent) and cleared from pending_title — the branch lands in the parent's lineage instead of falling back to a message-preview name. Persistence is best-effort: a broken DB logs and lets create succeed, leaving the lazy first-prompt path as fallback. Plain drafts keep the lazy-row contract unchanged. Fixes #93959 |
||
|
|
5fee992696 |
test(gateway): build-failure stub must fire its ready event
prompt.submit now retries a completed failed build (installing a fresh unset agent_ready before rebuilding), so a no-op _start_agent_build stub leaves the patient wait blocking forever — the turn thread outlived the test and the whole file hit the per-file timeout on CI. Stub a faithful failing build instead: set agent_error, fire the session's current ready event. The pinned contract (visible failure, no silent drop) is unchanged. |
||
|
|
67de93862c |
fix(tui-gateway): gate the ws-orphan interrupt of running turns on activity staleness
The 20s ws-orphan grace (
|
||
|
|
db339f0051 |
fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its own writer connection, self._lock, close-time WAL checkpoint, and token-writer thread. With N independent writers on one WAL file, one connection's close-time checkpoint could race another's growth — the lost/reordered-page-write signature across 11+ incidents (#90837). Adds hermes_state_registry.py: a process-wide, per-path, refcounted shared registry owning the writer boundary. - acquire(path): same resolved path returns the same instance (one writer connection, one lock, one token-writer thread) for every long-lived in-process caller (gateway runner, SessionStore, per-agent lazy recall, cron per-job, mirror, channel_directory, slash_commands, shutdown_flush, session_search, react_to_message, delegate, mcp_serve, auto_archive, tui_gateway). - close() on a shared instance is a NO-OP — the registry owns the lifecycle, so one caller's close can never tear down a writer other callers still hold. - Generation-aware retirement on inode change: a replaced state.db RETIRES the live generation (never lent again) but keeps it alive for existing holders; release is object-keyed so holders of the old generation drain it independently of the new one. The old generation's own write path still fails with the typed StateDbReplacedError (existing protection, unchanged). - Replacement-open failure leaves NO registry entry for the path — the next acquire retries fresh, never hands out a closed stale object. - All teardown runs OUTSIDE the registry lock: a final release's WAL checkpoint can never stall acquisition for every state.db. - close_shared_session_dbs() at gateway shutdown drains every generation (live + retired) as the final safety net. CLI one-shots, recovery flows, and read-only cross-profile opens keep using SessionDB() directly with their own close() — only long-lived in-process sites route through the registry. References #90837 (root-cause tracker stays open: the #10 EOF signature and the WAL-lifecycle A/B verdict remain under investigation there). |
||
|
|
a5f0fbb262 |
fix(sessions): per-session exclusivity is correctness, not a capacity policy
Cherry-picked from PR #94595 (author: Futahua) onto current main, with the maintainer-review revision points folded in during the rebase: - the lease engages UNCONDITIONALLY: try_acquire_active_session no longer returns a disabled no-op lease when max_concurrent_sessions is unset; the concurrency cap stays an orthogonal, optional policy checked second - ownership uncertainty fails CLOSED (SESSION_COORDINATION_UNAVAILABLE) instead of degrading to an untracked go-ahead: a corrupt/unreadable registry must not be collapsed into 'no owner exists' (review blocker 2) - the ownership admission sits at the _run_prompt_submit chokepoint that EVERY fresh turn source crosses, and crash auto-continue acquires (or bails) BEFORE emitting message.start — closing the #94778 bypass where backend B's auto-continue ran a duplicate turn while backend A was live (review blocker 1) - the TUI gateway claim helper fails closed on claim exceptions for every surface, not just desktop - CLI and messaging-gateway call sites pass live_session_id metadata so the (pid, live id) re-entrancy identity protects them from self-fencing on a leaked lease Co-authored-by: teknium1 <teknium1@users.noreply.github.com> |
||
|
|
4396253aed |
fix(gateway): gate compute-host interrupt forward on hosted activity
Follow-up to the salvaged #98571: forward the interrupt to the compute host whenever the parent 'running' mirror is stale, but only for sessions that actually have hosted activity — HostSupervisor.interrupt() calls start(), so an unconditional forward would spawn a compute-host child just to deliver an interrupt for an idle lazy session. Adds a regression test asserting the idle-lazy-session no-spawn path. Refs #92916 |
||
|
|
3ae74119c9 | fix(gateway): relay compute-host clarify state | ||
|
|
e17fd0a708 |
fix(state): decode errors now reach the heal path and fail loud in TUI (residual #98924 surfaces)
Companion to #98935, which fixes _fts_table_probe itself. This covers the surfaces that PR does not touch: - web_server._open_session_db_at_path: the one-writable-open heal only caught sqlite3.DatabaseError; a raw UnicodeDecodeError (pysqlite failing to decode SQLite's own error message over corrupt file bytes) bypassed it, so the heal documented for malformed schema never fired (#98924 Failure 1). Both catches widened; decode errors dispatch to the heal. - SessionSchemaMixin._recover_stale_fts_locked: drop-and-recreate skipped vtables whose probe raised UnicodeDecodeError, the same too-narrow catch the issue identified in the probe. - TUI gateway: _ensure_session_db_row returned silently when the store could not open, so prompt.submit streamed the turn while persisting nothing (#98924 Failure 2). It now returns False and prompt.submit fails the RPC with code 5072 so desktop maps it to a toast, mirroring the disk-full/5070 convention. session.create stays silent per its pinned degraded-mode contract. |
||
|
|
6874b99d49 |
fix(sessions): stamp launch-profile name on new session rows instead of NULL
Sessions created on the launch/default profile were persisted with profile_name = NULL by all three writers (run_agent._ensure_db_session None'd out 'default'; the desktop backend's _ensure_session_db_row and session.branch passed None when no profile_home override was set). NULL used to mean 'launch profile' by convention, but the desktop now keys sessions by (profile, id), filters the sidebar by profile scope, and resolves @session:<profile>/<id> deep links by profile match — a NULL row matches nothing, so sessions created around a profile switch vanished from the sidebar and their deep links could not be opened (#99222). The #94724 one-shot legacy-owner backfill stamps literal 'default' onto old NULL rows, so writers minting NEW NULL rows after that backfill ran recreated the exact state it exists to repair. Stamp the real profile name at creation time in all three writers. E2E-verified against a temp HERMES_HOME: both the desktop create path and the agent path now persist profile_name='default'. Fixes #99222 |
||
|
|
86a2fdc634 |
feat(tui): status rule shows cache-hit %, latency, t/s and honors display.status_bar.fields
Extends PR #98250's classic-CLI status-bar upgrades to the Ink TUI: - tui_gateway/server.py _get_usage() now emits cache_hit_pct, avg_latency_s, avg_tps (reads the same per-call deque history from agent/conversation_loop.py; keys omitted when no data — Codex app-server has no latency, zero cache reads show no %) - StatusRule renders the three read-outs as width-budgeted tail segments (breakpoints 96/104/110 cols, lowest priority — they shed first on narrow terminals) - display.status_bar.fields (the SAME key the classic CLI honors) filters TUI segments too: cache_hit, latency, tps, duration, compressions, bg_tasks, bg_subagents, voice, battery, title, context_pct, context_detail - values ride the existing usage payload/ticker; constants between events so the usage==last dedup keeps suppressing repaints - 3 new server tests, 5 new TUI tests; full ui-tui suite 1727 green |
||
|
|
4e860c093a | fix(tui): harden failed build recovery handoff | ||
|
|
258d0283b0 | fix(tui): close failed resume recovery races | ||
|
|
dbb9854591 | fix(tui): recover model switch after failed resume | ||
|
|
99a6852019 |
fix(tui-gateway): heal or fall back when a resumed session's provider is stale
A session row persists the provider identity a chat actually used. When that provider is later renamed or removed (e.g. a custom_providers:/providers: entry deleted, or a provider renamed oldone->newone), Desktop/TUI resume restores the stale name into agent init and dies with: agent init failed: Unknown provider '<name>' while the CLI resumes the same session fine with the configured default. - runtime_provider: add is_routable_provider() (full resolution chain: built-in -> providers: -> custom_providers: -> models.dev) - _stored_session_runtime_overrides: heal a non-routable provider via canonical_custom_identity (base_url -> model -> configured provider), drop to the configured default when unrecoverable, and clear the stale base_url after healing so a dead endpoint cannot override the registry URL - _start_agent_build: gate deferred-resume overrides on provider routability; when the stored provider is gone, prefer the model the user picked for THIS session, else the configured default - tests: is_routable_provider cases, heal/fallback round-trips, gate checks Refs #75128 |
||
|
|
b61408e95e |
feat(commands): attach desktop slash metadata to the registry
New commands and plugins declare argument_mode and desktop availability on CommandDef / register_command. commands.catalog ships that map so desktop does not need a second command list. |
||
|
|
c760143935 | fix(tui-gateway): claim disconnect sessions before teardown | ||
|
|
c19849cd02 |
fix(desktop): stop the status-stack poll storming a dead session with 4001s
The composer status stack polls `process.list` every 5s while a background
process row is on screen. `process.list` is session-scoped, so against a
runtime id the gateway no longer holds it returns 4001 "session not found".
`refreshBackgroundProcesses` swallowed *every* failure with a bare `catch {}`
commented "transient socket loss". A gone session is not transient: the poll
re-sent the same dead runtime id every 5 seconds for the lifetime of the
window. On one machine this produced 31,518 gateway rejections in a day
(vs 663 the day before), 18,614 of them against a single runtime id, and it
is what users see reported as "sessions stopped with a session not found
error" after an update.
The trigger is a reconnect, not the poll itself: anything that mints a fresh
runtime (gateway restart, the #94219 reconnect/replay work, an idle-reaped
pooled backend) strands the id the status stack is still holding, and nothing
in this path ever re-checked it.
Distinguish the two failure classes:
- 4001 / "session not found" is TERMINAL for that runtime id — latch the id
and stop polling it.
- A timeout or transport error is transient — keep retrying, since the
session may well still be alive. Misclassifying that direction would
silently freeze the status stack on a healthy session.
The latch is cleared when the status stack (re)binds a session id, so a
session that comes back under a fresh runtime resumes polling normally
rather than staying dark for the life of the app.
Also name the method in the gateway's 4001 warning. That line was added in
|