8099745a4e23cd07aa1ede41cc65cb0bb12a427f
33468 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8099745a4e |
fix(sms): scope TWILIO_PHONE_NUMBER per profile
__init__ and _standalone_send read TWILIO_PHONE_NUMBER via raw os.getenv, right next to the already-scope-aware TWILIO_ACCOUNT_SID/TWILIO_AUTH_TOKEN (_get_scoped_secret, established by merged PR #76664). Under gateway.multiplex_profiles, a secondary profile with its own Twilio account would silently send replies from the default profile's bridged TWILIO_PHONE_NUMBER instead -- Twilio rejects a from-number not owned by that profile's account (auth failure), or the message gets attributed to the wrong sender. __init__'s path is currently unreachable via normal secondary-profile adapter construction ("sms" is in gateway/config.py's PORT_BINDING_PLATFORM_VALUES, so _start_one_profile_adapters refuses to construct a port-binding platform for a secondary profile) -- fixed anyway for consistency with its sibling reads and to avoid a latent bug if that guard's scope ever changes. _standalone_send's path IS reachable: it's the generic cron/home-channel out-of-process delivery hook, which runs inside _profile_runtime_scope. Swaps both reads to the adapter's own existing _get_scoped_secret() helper, matching its sibling account_sid/auth_token reads exactly -- no new mechanism needed. Adds a TestMultiplexProfileScope test class to tests/gateway/test_sms.py mirroring the established Buzz/Discord/Telegram/WhatsApp/LINE/DingTalk/ Teams coverage for this bug class. (cherry picked from commit 2aeb3148367ecf87598d4458221a5fbf30c0af01) |
||
|
|
08830efd96 |
fix(secrets): secret-source re-pull no longer latches an empty snapshot or wipes sibling profiles
Symptom (#102041): under a multiplex gateway the default profile's vault/1Password/
Bitwarden/plugin-sourced credentials vanished for the rest of the process after the first
cron fire or the post-discovery plugin refresh; with the key already in the process env
(systemd EnvironmentFile=) the scope was empty from boot. Every get_secret() read then
failed closed ("No usable credentials", every Telegram sender rejected).
Why: _apply_external_secret_sources marked the home applied after any real fetch, but only
snapshotted names in report.provenance — the NEWLY applied ones. On a re-apply the previous
apply's own write-back makes every key `skipped_existing`, so the snapshot latched to {} and
_hydrate_profile_secret_sources returned that empty snapshot forever. Separately,
reset_secret_source_cache() was process-wide, so one profile's cron re-pull dropped every
sibling's hydrated snapshot (
|
||
|
|
580322ef1e |
fix(mcp): stdio MCP children get the routed profile's vault secrets, not the default's
Under a multiplexed gateway, `_build_safe_env` forwarded `os.environ[name]` for every name tagged in the process-global `_SECRET_SOURCES` map. That map is filled by EVERY served profile's secret-source hydration, while `os.environ` only ever holds the LAUNCH (default) profile's values — so once any profile's 1Password/Bitwarden source supplied e.g. GITHUB_TOKEN, every profile's stdio MCP server was started with the default profile's token. Resolve those names through the active profile's secret scope (`get_secret`) instead: the routed profile's value, or omitted when that profile has none. Under multiplex `get_secret` never falls through to environ; single-profile runs keep the .env overlay + environ behaviour, so the existing "vault vars reach MCP subprocesses" contract still holds there. `secret_source_names()` exposes the tagged NAMES only — values are never read from the shared map. Docs: the multi-profile guide's "MCP subprocesses only see their own profile's secrets" claim is now true for source-injected names too; say so explicitly. |
||
|
|
89fb1d028e |
fix(mcp): retain profile secret scope during discovery
(cherry picked from commit f344095fcf16c5d154ef2d207cce9b13339bf2e4) |
||
|
|
2952dc62bc |
fix(feishu): drive-comment turns and WS-thread callbacks run under their own profile, not the launch profile
Under `gateway.multiplex_profiles`, a Feishu adapter is built and connected inside `_profile_runtime_scope` (HERMES_HOME override + secret scope as contextvars), but two hops started from an EMPTY context and so executed under the LAUNCH profile: - `feishu_comment.handle_drive_comment_event` ran the whole comment AIAgent turn on a bare `loop.run_in_executor(None, ...)`: model/credential resolution raised `UnscopedSecretError` (silent empty reply), or — when the default profile held the same key — used the default profile's config/model/state for a secondary profile's doc. - `FeishuAdapter._connect_websocket` ran the lark WS client on a bare executor thread. The SDK fires every event/card callback on that thread and they hop back to the adapter loop via `run_coroutine_threadsafe`, which copies the CALLER's context — so all pre-handler work (inbound media caching, `.update_response` marker, FEISHU_REACTIONS env, drive comments) ran unscoped. The pending-inbound drainer thread spawned from that callback had the same shape. Carry the scope across each hop with `contextvars.copy_context().run`. For the SDK-owned thread the snapshot is taken once at `_connect_websocket` (inside the profile scope; the restart supervisor task inherits it too), so no per-callback re-scoping is needed. Supersedes the `_submit_on_loop` re-scoping approach of #63962 (nateEc, earliest fix): scoping the WS thread at its source covers every callback without rebuilding the secret scope per call. Co-authored-by: Nathan Shan <nathanielcrush51@gmail.com> |
||
|
|
2b4deeb32b |
fix(auth): key sibling per-process credential memos by profile home under multiplex
Same class as the resolve_nous_access_token memo: three more process-wide memos carried a credential resolved under one profile's HERMES_HOME override into another profile's turn for their TTL. - hermes_cli/nous_billing.py::_token_cache (30s (token, base) memo for the charge poll loop) was a single unkeyed slot -> dict keyed by hermes_home_key(); invalidate_cached_token() clears the dict. - agent/moa_loop.py::_runtime_cache carried api_key/base_url/api_mode keyed only (provider, model) for 5 min -> (hermes_home_key(), provider, model). - agent/auxiliary_client.py::_client_cache_key had no profile component, so callers that omit api_key (pool / Nous auth.json paths) could be handed a client built with another profile's bearer -> hermes_home_key() leads the key. WHY hermes_home_key(): it reads the per-turn HERMES_HOME override the multiplex gateway sets (falling back to the env var), and it is symlink-stable, so the memo key is exactly the credential home the resolution itself read from. Profiles stay independent islands; the default-profile process env never leaks into a secondary's turn. Tests: one invariant per site, proven red on origin/main. |
||
|
|
173105ce6f |
fix(auth): scope the resolve_nous_access_token memo to the active profile
resolve_nous_access_token()'s 5s startup-burst memo (#76930) cached the resolved Nous Portal access token in a single module-level slot keyed by nothing but wall-clock time. The underlying resolution is profile-scoped: _auth_file_path() reads get_hermes_home(), which checks the context-local _HERMES_HOME_OVERRIDE ContextVar before falling back to the HERMES_HOME env var — gateway/run.py and tui_gateway/server.py set that override per-profile for multiplex concurrency. In a multiplex gateway serving two profiles with different authenticated Nous accounts, if profile A's context resolves a token and profile B's context calls resolve_nous_access_token() within the next 5 seconds, profile B received profile A's cached token — used to authenticate against the managed tool gateway / relay self-provisioning under the wrong account. Key the memo by str(get_hermes_home()) instead of a single slot, so each profile's context reads only its own cached token. The lock around read/write is unchanged; only the cache's shape moved from a single (timestamp, token) tuple to a dict keyed by resolved home. (cherry picked from commit 9d8846b88cfd2ea74c6958d5f8f28a50880dda75) |
||
|
|
3b044261b6 |
fix(gateway): media denylist covers every profile's credentials, not just the launch home
Under `gateway.multiplex_profiles` one process serves every `<root>/profiles/*`,
but `_media_delivery_denied_paths` expanded `_ROOT_CREDENTIAL_PATHS` only under
the import-time `_HERMES_HOME` / `_HERMES_ROOT`. A `MEDIA:<root>/profiles/<B>/.env`
(or auth.json, state.db, config.yaml, sessions/, mcp-tokens/) emitted in ANY
profile's turn — including B's own, whose HERMES_HOME override was never
consulted — passed validation and was natively uploaded to the chat. The ALLOW
side (`_profile_cache_roots`) already enumerated profiles at check time; the
DENY side did not.
`_credential_home_roots()` now yields the active `get_hermes_home()`, the shared
root and every `<root>/profiles/*` at check time (shared `_profile_dirs()` with
the allow side), and the denylist is built from that. Profile cache artifacts
and plain agent-written files under a profile stay deliverable.
Live repro (/tmp/mux_audit/fix-media-denylist/repro.py): before, all five of
profiles/B/{.env,auth.json,state.db,config.yaml,sessions/s1.json} validated as
deliverable while <root>/.env was blocked; after, all None, cache/images/gen.png
and report.pdf still deliverable.
No prior report. Write-side analogue: #107327 / #107335 (memo keying, different
mechanism — left as is).
|
||
|
|
942973ae90 |
fix: every Bedrock client rebuild lands on the startup wire (Claude SDK, Converse region, guardrails)
Follow-up to
|
||
|
|
6c3d4a4af7 |
fix(desktop): resolve plugin SDK namespaces lazily, not at module scope
runtime.ts captured the SDK namespaces (the plugin SDK, React and the two jsx runtimes) in a module-scope object literal, and that module sits in an import cycle: sdk/index -> @/contrib/* -> contrib/runtime-loader -> sdk/runtime -> sdk/index. In an unbundled (dev) graph the namespace object is live, so the capture works. In a production bundle the bundler emits the SDK namespace as a hoisted `var` whose assignment lands AFTER the literal that reads it, so the captured value is `undefined` (no TDZ error) and `Object.keys(GLOBALS[globalKey])` in `shimUrl()` throws "Cannot convert undefined or null to object". That throw happens inside `loadRuntimePlugin()` -- via `unsupportedImports()` -> `sdkImportMap()` -> `shimUrl()`, which run for every source before it is evaluated -- so it is content-independent: EVERY plugin loaded from $HERMES_HOME/desktop-plugins/<id>/plugin.js (and the desktop/plugin.js half of a unified package) fails to load, showing status "failed" in Capabilities -> Plugins. Resolve the namespaces at call time instead: `installPluginSdk()` and the shim builder only ever run once the app is up, so reading them there is always safe and statement ordering can no longer matter. Repro (production build only): build apps/desktop for production, drop any plugin.js into ~/.hermes/desktop-plugins/<id>/, start the app -> [plugins] runtime load failed (<id>) TypeError: Cannot convert undefined or null to object (.../assets/sdk-<hash>.js:5). Note: a vitest unit test cannot catch this (dev module graph keeps the namespace live); the faithful guard is a production-build smoke test that loads a fixture plugin through loadRuntimePlugin(). |
||
|
|
4bdd64b334 |
The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an anonymous account exactly one model on its own host, refuses everything else with a structured 429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and tells a signed-in account that still asks for `nous/welcome` what to switch to in an `x-nous-model-switch` header. Four client-side gaps against that contract: - Auxiliary calls were refused on every session. The auxiliary client asked the welcome host for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free` before each fallback. On the welcome host it now uses `nous/welcome` (its backing model covers auxiliary work) and skips Nous for vision, which the welcome model does not take. - The structured 429 body was never read. The classifier now parses `reason` / `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited` are rate limits that honour `retry_after` and never rotate the free tier's only credential. The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and fall back instead of retrying or re-exchanging. The terminal paths say what happened and name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal). - The `x-nous-model-switch` header was ignored. The chat-completions transport records it beside the rate-limit and credits headers; the next call moves the session, and the config default when it still names `nous/welcome`, to the backing model the gateway named. - A guest fell back to the paid host. With `inference_base_url` absent from the exchange or outside the host allowlist, routing defaulted to inference-api, where every request is a 400. A guest now defaults to the welcome literal at the exchange, in the shared store's shape, and in effective routing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8) * feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime rung now sets it up there instead of failing "not logged in", so the guided chat no longer races the root profile's first-run mint. nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever nothing else is configured; "on-request" mints only when the free tier is asked for by name (nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers — the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the connector token path — still adopt what the shared store holds, so every profile follows the one identity the guided setup created, but never create one on their own. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c) (cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165) * feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity (nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is configured. "explicit" means Hermes never creates one on its own: the only creator is the new provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and before the guided chat exists — so the identity lands in the root store every profile reads through and is there before any session asks for nous/welcome. That closes the race against the backend's own setup, and makes "only when the setup-bot flow is used" literally true. The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome (the hermes model row, --provider nous), which treated a model name as intent and was broader than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to sign in from. Implicit callers still adopt an identity the shared store holds, and a retired credential is replaced (a continuation, not a creation). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0) * fix(auth): remove the nous.guest_setup knob; the free tier is created on first use `nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`, unknown values read as `auto`, and under `explicit` a CLI-only install could never get an identity, which contradicts the first-run contract (first command mints, then chats). The mint race the knob accompanied is already benign: every caller takes the profile lock then the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which stays. `nous.guest` remains the only free-tier policy. Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through `ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config entries. The three policy tests that hold regardless of the knob are kept under `TestExplicitProvision`; the two that only tested the knob are deleted. (cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261) * fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch` moves the live session to that model and moves `config.yaml`'s default off the alias in the same step. The messaging gateway's fallback-eviction check compares the agent's model with the config default and evicts on any mismatch that is not a /model override, so when the config write did not land (unreadable config, lock) the cached agent was evicted once per turn, and prompt caching with it. `apply_model_switch` now stamps the alias it moved the session off on the agent, and `_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as deliberate, beside the existing /model override case. The check takes the agent and the config model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them. (cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954) * fix(auth): the free tier outranks implicit host credentials in provider resolution On a fresh install with a leftover ~/.aws profile, resolve_provider("auto") reached the Bedrock rung before the free-tier rung, so the first turn ran on Bedrock and failed 403 while the free tier was still being minted in the background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s, three retries, no answer; the next process then switched to nous/welcome. The free-tier rung now sits directly above the Bedrock chain: when nous.guest is on, an existing free-tier identity answers, else a blocking mint runs, and only then does the boto chain get a say. Everything above is unchanged and still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter pool, a logged-in active_provider. nous.guest: false skips the rung, and a failed mint still falls through to Bedrock and the no-provider guidance. Tests: six precedence cases (identity present, fresh mint, free tier off, env key still wins, sign-in still wins, failed mint falls through). The opt-out test now neutralizes the AWS chain like the precedence tests do; on a machine with ~/.aws it was failing for the same reason as the bug. Live after the fix, same Mac, AWS credentials visible, isolated shared store: identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in 11 s. (cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a) * fix(auth): review follow-ups for the free-tier rung (NS-829) - tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches the free tier off; its contract is the boto chain, and the free tier now sits above it. - gateway/run_notifications.py: the free-tier startup line reads auth.json before consulting the resolver, so a gateway boot on a machine with AWS credentials never mints or refreshes over the network. - hermes_cli/anon_auth.py: module docstring says where the free tier sits in the ladder instead of "the ladder is untouched". - tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of six (parametrized ladder cases; a failed mint that returns None or raises falls through to Bedrock). scripts/run_tests.sh on the five affected files: 147 passed, 0 failed. (cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11) * feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone The free tier is pre-GA. Until GA it must not exist for anyone who did not ask for it: no identity minted, no portal traffic, no free-tier copy on any surface. One environment variable now decides that, and one function reads it. `guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly "1"; only then does `nous.guest` (the user's off switch) get consulted. Every free-tier site already funnels through `guest_enabled()`, so the gate closes minting, routing, connector entitlement, status lines and the picker row in one place. With the variable unset, `resolve_provider("auto")` on a fresh install raises `no_provider_configured` exactly as upstream does. `HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted identities as a side effect of provider resolution, and `_has_any_provider_ configured` read them ahead of every other check, making the CLI a second reader of a flag that must have exactly one. `_forced_new_done` and the `force` parameter of `_reconcile_and_provision` go with them. Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847). Not a user preference: the variable is never written to config.yaml or .env and never shown in setup. It is deleted at GA together with its comment in anon_auth.py. This is a deliberate, temporary exception to the "no new HERMES_* env vars for non-secret config" rule. Tests: fixtures set the gate instead of deleting the old lever; one new invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "", "0", "true" and "new" all leave the tier off with zero portal calls, red on the previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the feature. * feat(auth): the free-tier identity is created in one place, at boot; every other site is a read Before this commit eight sites could create a Nous free-tier identity as a side effect of something else: resolving a provider, the CLI's first-run check, the CLI's session setup (in the background beside an own key), a connector bearer read, the desktop polling `free_tier.status`, the sign-in precondition, the desktop's `free_tier.provision`, and the dead-credential re-mint. A poll could mint. Provider resolution could hit the network. Two of them raced each other on a fresh install. Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator. `hermes serve` runs it on a daemon thread from `_lifespan` beside the other background boots; `cmd_chat` runs it synchronously before the first-run guard. It inventories credentials first (`resolve_provider("auto", skip_free_tier=True)`: what would carry inference if the free tier did not exist), creates the identity only when `guest_enabled()`, resolves inference, records a `SetupRecord` in process memory and broadcasts ONE `setup.ready` event. It runs on every boot; only the mint is gated. `ensure_portal_identity` now requires `explicit=True` and raises otherwise. Its callers are the bootstrap, the desktop's `free_tier.provision` (the explicit retry when the boot could not create the identity) and the two dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`, `managed_tool_gateway._replace_dead_guest_token`). The background thread path and `provision_free_tier` are deleted with their last callers. Reads that used to mint and now only read: `auth.py::resolve_provider` rung 7 (an existing identity still outranks the Bedrock chain, NS-829 ordering kept), `main.py::_has_any_provider_configured`, `cli_agent_setup_mixin._ensure_runtime_credentials`, `managed_tool_gateway.read_nous_access_token` (no identity -> None), `anon_sign_in.run_sign_in` (no identity -> Unavailable), `methods_free_tier` `free_tier.status`. `setup.status` answers from the record for the launch profile, blocking up to 8 s while the bootstrap is in flight so a client's first poll lands after the identity exists rather than racing it; a named profile, or a process that never ran the bootstrap, keeps today's live probe. The record's fields ride along additively (`ready`, `free_tier`, `other_providers`, `inference_provider`). Identity and inference are decoupled (NS-845 Q1.3): the mint sets `active_provider="nous"` only when the inventory found nothing else usable (`_mint_locked(carries_inference=)`); an adopted account always does. A token refresh no longer re-elects the provider it refreshed (`_save_provider_state_to_source` writes credentials, not the user's choice) — that write was how an own-key install ended up on the free tier after the first connector call. Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check), bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint), 62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and `provision_free_tier`), and a04b05260c (blocking mint in the resolver). Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847. Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps inference; reads never reach the portal; a refused mint is memoised), `free_tier.status` fails loudly if it ever calls the creator, the resolver stub fails loudly if resolution ever mints, `setup.status` reads the record, `skip_free_tier` proves the inventory question. The three sign-in tests for the deleted pre-mint collapse into one (`no identity -> Unavailable, zero portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20. * fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup" A free-tier identity carries $0 by design, so the portal seed reports `paid_access=False` for it. `is_free_tier_model` did not know the welcome host, read that as a depleted account, and every free-tier turn ended with the credits-depleted notice telling the user to top up an account they do not have. Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host (`anon_auth.route_is_welcome_host`) is the free tier. The host is the evidence, not the model name: the paid inference host can serve `nous/welcome` to a named account and that account's depletion is real, so `("nous/welcome", <inference host>)` stays False. Local data only, like the three rules above it. Restores the two contracts dropped by hermes-magic 674e11d1eaa (the prototype line ran without unit tests): the welcome host is free without any pricing evidence; the model name alone is not. The first is red without this fix. * fix(copy): free-tier text stops promising a connector transfer and never names the config key Sign-in copy on every surface said "Sign in to keep your connectors" and ended with "Your connectors are kept." The transfer registry that would make that true is empty (NS-821): nothing carries over today. The copy now says what signing in does give ("unlock more models and tools") and the completion line names the account, not a transfer. The docs page loses the "connectors carry over" paragraph for the same reason. The picker's off-state line exposed `nous.guest: false` and the word "guest"; user copy names the free tier only (R-USR-1). The docs page gains the pre-rollout note: until GA nothing on it happens without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and "replaced on next use" sentences now describe the boot bootstrap. zh is a strict locale: the `freeTier` block was English placeholder text copied from `en`; it is now Chinese. `connectorsKept` is renamed `completedBody` since it no longer talks about connectors. * feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1" as on. Until now nothing in the desktop set it, so a packaged app could never turn the free tier on, and a backend spawned by the app could disagree with the app about whether the tier was live. `electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled` is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has `--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch into a module constant. `desktopBackendSpawnEnv` wraps every backend env as the outermost call and writes the flag LAST, as "1" or an explicit "0", so no earlier spread (`process.env`, `backend.env`) can resurrect a stray value from the parent shell. Stamped onto all three spawn sites: the primary `serve` spawn, the pooled per-profile spawn, and the remote SSH `exec env ...` command (which gains ` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and the backend probes are not backend spawns and do not get it: a `hermes --tui` typed in the pane must not mint. The renderer learns the same fact read-only through the existing `hermes:launch-flags` sync IPC (`guestOnboarding`) and preload (`window.hermesDesktop.guestOnboardingEnabled`). Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding` maps to it in main). Two invariant tests on the pure helpers: only "1" or the argv flag enables; the spawn env carries "1"/"0" as the last word and preserves every other key. * feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll The backend's boot bootstrap now announces `setup.ready` once, after it has created (or refused) the free-tier identity and resolved the inference route. The renderer used to discover both by polling `setup.status`, `setup.runtime_check` and `free_tier.status` every 60 s from `useStatusSnapshot`; a fresh install's chip, notice strip and onboarding overlay could sit stale for up to a minute after boot, and three RPCs a minute per window kept asking a question whose answer changes only at boundaries the backend already announces. `handleLifecycleEvent` routes `setup.ready` (active source only, like `skin.changed`) to `notifySetupReady()`, a one-shot tick atom in `live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens to it and runs one readiness round at once (`setup.status` + `setup.runtime_check` + `free_tier.status`). The readiness legs also run once on open and on return from another app, as today. The 60 s tick keeps only `getStatus()`. `SetupStatusSnapshot` types the record's additive fields (`ready`, `free_tier`, `other_providers`, `inference_provider`); readiness semantics are unchanged and still key on `provider_configured` + `runtime_check`. Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one refresh from the active source and none from another; the snapshot hook's contract is three legs on open, one leg on the tick. * fix(cli): the banner names the free tier's model instead of "no model configured" The welcome banner prints before credentials resolve, so on a fresh install `model` is empty and the banner said, in red, "no model configured — run /model or hermes setup". Under the free tier that is false: the route is already known from local state (identity on disk, tier on), and the first message will run on `nous/welcome`. `_banner_left_lines` now asks the route the same question when `model` is empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the banner's 'no model configured' line reads the resolved route"). Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`; gate off -> the red line, zero portal calls. * fix(aux): vision on the free tier uses nous/welcome too The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the welcome host would have sent every image step past the free tier for no reason, so the auxiliary client pins the route's one model for every lane. A backing model that takes no images answers with the upstream's own error, which the ladder handles as it always has. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415) * fix(gateway): hermes gateway run is a boot owner of the free tier too Rung 5 made every demand-time free-tier site a read: resolve_provider, the connector token, the /login precondition. That is only correct if every process that can reach those sites ran the bootstrap first. The CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone messaging gateway did not. A fresh HERMES_HOME with the gate on and `hermes gateway run` reached provider resolution with no identity to consume, and /login returned Unavailable. Reported by @andrexibiza on #107697 (P1). GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an executor thread right after startup recovery and BEFORE any adapter connects, so a fast first DM cannot arrive with nothing to resolve. It is its own step, not part of the turn-machinery warm-up: the warm-up is an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0); the bootstrap is correctness and must always run. With the gate unset it is a local inventory and no network. Live, real GatewayRunner.start against a fake portal in a fresh home: gate on -> 1 create, identity persisted, resolve_runtime_provider=nous, /login precondition sees the identity gate off -> 0 portal calls, no identity, no_provider_configured Before the fix the gate-on row was identical to the gate-off row. Test: the bootstrap seam runs before _start_prefilter_platforms and delegates to the one creator. Red on 5554eb6993 (no seam), green here. --------- Co-authored-by: Robin Fernandes <robin@soal.org> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
cbcf7b72f7 |
feat(gateway): sign in with a Nous account from a chat (/login), one shared sign-in flow (#105261)
* refactor(auth): one sign-in flow behind SignInState, rendered by the CLI and the desktop * feat(gateway): /signin signs the free tier into a Nous account from a DM * feat(cli): chat surfaces name /signin as the sign-in verb * fix(auth): review follow-ups for the shared sign-in flow and /signin * fix(i18n): carry the /status free-tier line in every locale catalog * refactor(cli): the chat sign-in command is /login * fix(auth): durable override cleanup in the /login sweep, and the sign-in flow in its own modules |
||
|
|
3b01b4ce0f |
feat(desktop): Nous free tier on Hermes Desktop (#105260)
* feat(desktop): free-tier state over RPC, status routes that name it, and a sign-in that keeps connectors The desktop learns about the Nous free tier by reading local auth state (pull): free_tier.status answers has_guest / enabled / carries_inference / notice_pending with zero network, and free_tier.ack_notice persists the one-time notice flag on the identity itself. setup.runtime_check reports free_tier for the selected route; /api/portal, the Nous card in /api/providers/oauth and billing.state carry free_tier (billing answers the free tier locally instead of a portal call that can only fail). The free-tier picker row carries an explicit free_tier_row flag and is never priced or locked. POST /api/providers/oauth/nous/start over a free-tier identity registers the connector transfer and returns its code and consent URL; the poller waits for the transfer before the token grant, persists the account, runs settle_after_upgrade, and the poll response gains reason, account_email and model. * feat(desktop): free tier on Hermes Desktop: ready screen, notice strip, status chip, Billing view, one sign-in dialog The renderer reads the free tier from free_tier.status (pull) into one store; the first-launch intro is the same state rendered two ways, keyed on the backend's one-time flag: the onboarding overlay opens on a ready screen when the free tier carries inference, else a one-time strip above the composer. Settings > Billing gains a free_tier view (notice with one Sign in, Plan / Model / Connectors summary, plan card, footnote; no payment or usage rows). A status-bar chip names the tier and model while it carries inference. Every entry point opens one claimed sign-in dialog that drives the extended oauth/nous route and maps the poll's status and reason to the ruled screens; Done settles billing, model options, providers and re-homes a session still on nous/welcome. The picker badge also fires on free_tier_row. Docs: Desktop section in the free-tier guide, AGENTS notes. * fix(desktop): free_tier.status starts the free tier's background setup when no identity exists A served backend has no session-setup moment like the CLI's, so beside an explicit provider the free tier was never set up on the desktop: no connectors, no notice strip. The first status read now starts the same one-attempt background setup; the call itself never waits. * fix(desktop): one Sign in on the Billing page; Settings > Providers names the free tier, never Connected The free-tier plan card is the what-you-get text alone (the notice carries the page's one Sign in). The Nous provider row reads Nous · free tier with a Free tier tag while the identity is the free tier, instead of Nous Portal · Connected. * fix(desktop): Settings > Providers never files the free tier under Connected * fix(desktop): the intro's shape is keyed on the route, not on the identity free_tier.status reports available (an identity exists and the tier is on); whether inference runs on the free tier is setup.runtime_check.free_tier, keyed on the resolved endpoint. The ready screen shows when that route is the free tier; the composer strip when the user's own provider carries inference. An own-key install used to get the ready screen. * docs(desktop): say what the free-tier chip is keyed on * fix(desktop): the featured Nous row's pitch on the free tier says what signing in adds * fix(desktop): a cancelled or superseded sign-in attempt can no longer change the identity or hide the intro Four lifecycle holes from review. The Nous poller checks the session's cancelled flag after the transfer wait, after the token grant, and once more under the session lock together with the save, so a sign-in the user abandoned never persists. The renderer's sign-in store carries an attempt generation that every continuation checks after each await, so a poll from a closed attempt cannot publish over the one on screen (and its backend session is cancelled). The ready screen comes down only after the backend recorded the acknowledgement. A composer still mounted takes over the notice claim when its owner unmounts. One thin test per hole. |
||
|
|
04a76c4109 |
fix(auth): a sign-in from the free tier settles the default model and route (#105259)
* fix(auth): a sign-in from the free tier settles the default model and route once, for every caller Picking the free-tier row leaves model.default at nous/welcome pinned to the welcome host. After a sign-in an account cannot keep either: the welcome host refuses account tokens, and the portal host serves nous/welcome as a paid model. One completion step, anon_auth.settle_after_upgrade, now runs after the account is persisted: a config on the free tier's route moves to the account's inference host and the recommended default for the account's plan, through the same config write a plain Nous login uses; a config on the user's own model is left alone. The pick is the one GET /api/model/recommended-default already makes, factored into models.recommended_nous_default_model so the CLI and the desktop land on the same model. hermes auth upgrade prints the new default. * fix(auth): a sign-in completion with no eligible recommendation leaves no default model The static provider-wide default is not narrowed by the account's plan or org policy, so writing it as a fallback could persist a model the account may not use. When the recommendation cannot yield a model, the route still moves to the account's host but model.default is left unset; the CLI says so and points at `hermes model`. * docs(free-tier): say what happens when no recommendation is available after sign-in * fix(auth): sign-in completion moves the host and clears the default in one config write Two writes could fail between them and leave the account host paired with nous/welcome. _update_config_for_provider gains clear_default so the caller with no model to offer removes model.default in the same atomic write that sets the host. |
||
|
|
a2db110ccc |
feat(auth): Nous free tier: free inference and connectors out of the box, one command to sign in (#105258)
* feat(auth): Nous free tier core: anonymous identity minted on first use, welcome inference, shared-store scoping A fresh install with no provider sets up a free Nous identity (anonymous auth method of the nous provider) instead of forcing the setup wizard. The identity is persisted through the same path a real login uses, so the resolver ladder is unchanged. Two seams differ: token acquisition re-exchanges the anon credential (no refresh token), and routing pins the welcome host's single model nous/welcome. One identity per shared store; nous.guest: false turns the free tier off. * test(auth): free tier core contracts: lifecycle, resolver precedence, exchange seam, model pin * docs(user-guide): free tier and signing in New page explaining what a fresh install gets before any key or sign-in (free inference on nous/welcome plus connectors), how the free tier coexists with a user's own API key, how to sign in with hermes auth upgrade and keep connectors, how to turn the free tier off with nous.guest, what hermes logout does in each state, a troubleshooting table, and a plain privacy note. Wired into the Using Hermes sidebar. * fix(auth): logout leaves the free tier alone and clears the shared store for a real Nous account Logging out of the free tier is a no-op: it is not a login, so nothing is cleared and the user is told they were never signed in. Logging out of a real Nous account now also clears the cross-profile store, so a profile logout is not silently re-adopted on the next boot. * fix(model): switching off the free tier points at signing in, never hops providers * Name the free tier in the gateway startup notice and tell explicit-provider installs about it once * Render the Nous free tier as free tier on auth status, auth list, hermes status and portal info, short-circuit billing copy for it, and skip the keepalive when there is no refresh token * fix(auth): free tier is set up where nothing is configured: resolver last rung and first-run check Both the provider resolver's terminal rung and the CLI first-run check now try to set up the free tier before declaring nothing configured. On a fresh install the first command lands in chat on nous/welcome; a failed setup still falls through to the existing guidance. * Add hermes auth upgrade: sign the free tier into a Nous account while keeping its connectors The device-code flow runs as usual, with a promotion intent registered on the portal between the code request and the token poll so the account that approves the code inherits the free tier's connectors. The promotion status decides the outcome: only a completed one is followed by the token grant, which is persisted over the free-tier singleton and the shared store. Declined, superseded, retired and busy outcomes each print their own plain copy, and a retired identity is cleared so the next use sets up a fresh one. User-facing text never names the free tier's internals. * Show the Nous free tier as one picker row with nous/welcome and hide it when nous.guest is off * fix(auth): upgrade opens the consent page for this sign-in; one mint attempt per process; forced free tier wins the first-run check The browser leg of hermes auth upgrade now prints and opens the promotion claim URL with the claim code, not the generic device page. A failed mint is attempted once per process so several bootstrap sites cannot hit a closed gate or a 429 twice; a retired credential resets that so re-minting still happens. HERMES_FORCE_GUEST is honoured ahead of the first-run provider check. * fix(auth): pin the welcome model on the selected route, not on profile state; background setup retries after a failure A credential-pool entry can select a paid Nous key while the profile singleton is still the free tier. The model pin now keys on the resolved endpoint (welcome host) in agent init and /model, and the pin in model normalization is removed since it had no route to look at. A failed background identity setup releases its latch so a later attempt in the same process can try again. * fix(auth): decide the Nous model together with the route on every credential-pool swap The credential pool can move a Nous agent between the welcome host and the portal host after init. One helper, pin_model_for_route, now runs at init and inside every pool swap, so the welcome host always carries nous/welcome and a paid endpoint always keeps the caller's model. * fix(auth): apply the route model policy on every wire mode during a pool swap; release the setup latch if the thread cannot start * fix(auth): free-tier lifecycle takes profile then shared lock, reconciles with the shared store, persists the mint before exchanging, and clears only the identity that died The shared store is the identity of record for a Hermes root: a profile holding a stale free-tier identity adopts a sibling's newer sign-in instead of keeping the guest, and never overwrites the shared account. Locks are taken in the documented order (profile, then shared). A minted credential is stored as soon as create succeeds, so a rate-limited or timed-out exchange does not lose it and trigger a second mint. Retiring a dead credential removes only that credential from both stores. Guest exchange uses the resolver's canonical portal URL. * fix(auth): a credential rotation never rewrites the conversation model; connectors honour the off switch and replace a retired free-tier credential The welcome host serves one model, so a rotation onto it is refused for any conversation on another model instead of silently switching that conversation to nous/welcome (the model pin applies only when a route is first chosen). The connector token path now treats the free tier as absent when nous.guest is false, including cached tokens, and shares the one dead-credential rule with inference: a retired identity is replaced once rather than returning its stale token. * fix(auth): plain login never imports the free tier as OAuth credentials; the gateway startup line reads persisted state only A free-tier identity in the shared store is not an OAuth credential to offer for import; a real sign-in replaces it. The gateway's startup notice now answers provider precedence from persisted state (no token refresh at boot), so an expired free-tier token cannot stall the online message. |
||
|
|
2ddeba9e17 |
Merge pull request #99523 from somewheresy/justin/e-1047-route-hermes-actual-provider-through-chat-completions-with
fix(providers): use chat completions for all Actual routes |
||
|
|
564aef2946 |
fix: /model onto Bedrock Mantle keeps SigV4 auth instead of 401ing
Startup (agent_init._init_openai_client) ran configure_bedrock_openai_client_kwargs,
so the aws-sdk sentinel became a SigV4-signing httpx client. Every later client
rebuild — switch_model, fallback restore, credential rotation, request-scoped
clients — went through create_openai_client with bare {api_key, base_url} kwargs,
so the OpenAI SDK sent "Authorization: Bearer aws-sdk" and Mantle answered 401
"Invalid bearer token". Symptom: `hermes --provider bedrock --model
openai.gpt-5.6-terra` works, `/model openai.gpt-5.6-terra` inside the CLI/TUI fails.
Install SigV4 in create_openai_client itself (the single chokepoint every primary
OpenAI-wire client passes through) whenever the base_url is a Mantle host, so all
rebuild paths inherit the fix rather than each remembering to call the adapter.
A real AWS_BEARER_TOKEN_BEDROCK key is left alone (the adapter only rewrites the
aws-sdk / no-key-required placeholders).
Live repro (local HTTP sink capturing the Authorization header after switch_model):
before "Bearer aws-sdk", after "AWS4-HMAC-SHA256 Credential=...".
|
||
|
|
8a6b5b67a7 | fix(providers): block Actual at Responses send sites (E-1047) | ||
|
|
4135933ec9 | fix(providers): preserve Actual routing during startup and key reload (E-1047) | ||
|
|
8b2b83a906 | fix(providers): pin all Actual routes to chat completions (E-1047) | ||
|
|
d7b0a72c2a | fix(providers): route Actual through chat completions (E-1047) | ||
|
|
872bafd58d |
Merge pull request #107576 from NousResearch/fix/107109-probe-settle
fix(desktop): settle the login-shell PATH probe and kill the probe child on timeout |
||
|
|
f98cb00a8b |
fix(vault): 2FA review follow-ups — per-digit fill only for an unmistakable maxlength=1 widget; honour otpauth digits/period/algorithm
Reviewer findings on #107585: - build_otp_fills split any >=4 code-like controls into digits. A page with promo/zip/referral 'code' inputs next to the real OTP box would have had a digit sprayed across unrelated fields. Split now requires exactly len(code) controls that are all maxlength=1, same form, adjacent in DOM order (inspection JS exports maxLength); anything else fills ONE field, the best-scoring one. Verified on the real Browser Use stack: 6-box widget gets one digit each; scattered page fills only the one-time-code input. - normalize_otp_secret dropped digits/period/algorithm from otpauth:// URIs, so an 8-digit or SHA-256 authenticator would get wrong codes. Non-default parameters are now stored as seed|digits|period|algo and honoured (RFC 6238 SHA-256 8-digit vector added); hotp:// is rejected explicitly. - rebase on main (prompts.ts conflict) + prettier. |
||
|
|
d9ca9c974d |
feat(vault): two-factor codes — automatic from a saved authenticator key, otherwise asked for in the user's UI
Follow-up to #106480. Sites that ask for a code after the password stopped the agent cold: the login classifier excludes one-time-code fields on purpose (a password must never land in an OTP box) and there was no tool for the second step, so the only move was to ask in chat. browser_vault_enter_code Fills the one-time code the current page asks for. Two sources, same invariant as passwords (the code goes to the page over the supervisor socket and never enters model context): - a TOTP seed on the login: local vault `otp_secret` (RFC 6238, stdlib, verified against the RFC test vectors), 1Password `op item get --otp`, Bitwarden `bw get totp`. Nobody is asked. - no seed: the surface prompts "Verification code for {site}"; the user types what their phone/email/app shows. Enter on empty / Skip declines and the tool returns code_declined ("do not ask again this turn"). no_code_field tells the model the site wants a passkey / hardware key / app approval: hand it to the user's device and wait for navigation. Per-digit OTP boxes (maxlength=1 pattern) get one digit each in DOM order. Surfaces CLI: sudo-style panel, code shown as typed (not a secret worth masking, typos must be visible), Enter submits, ESC/empty skips. Desktop: "Verification code for {site}" card via vault.code.request / vault.code.respond (gateway), owner-routed like the other vault prompts. Settings → Passwords & Logins: optional "Authenticator key" field on the add form (base32 or otpauth:// link); items with one show a "2FA auto" badge. `hermes vault add` asks for the same optional key. browser_vault_fill's result now says what to do next ("if the site asks for a verification code, call browser_vault_enter_code with this handle"). Six locales. Verified live (real model, local 2FA site that checks the TOTP; CLI PTY): A. login saved with authenticator key → signed in through 2FA, zero prompts, code/password absent from the transcript B. login without key → code panel → user types code → signed in C. panel dismissed → agent stops and explains, never asks in chat Unit: RFC 6238 vectors, seed normalisation, mint-without-asking, per-digit spread, decline, no-code-field; Desktop card test (owner routing, trim, Skip). |
||
|
|
77e55b4d1f |
fix(desktop): show each in-app tip once, never lap the catalog again
The idle tip rotation walked the catalog as a ring: after the last tip it wrapped to the first, so a user who had already seen every tip kept getting "Start fresh", "Teach it once", ... again every six hours for as long as they used the app. Only the X stopped a tip, and letting a bubble time out (the normal way it leaves) counted for nothing. The walk is now one lap. nextTip also steps over every tip in the seen ledger ($tipShownAt, which already recorded every catalog tip that reached the screen), so a tip shows once however it left, and the rotation runs dry once every tip has had its moment. Settings > Reset clears the seen ledger and the cursor as well as the retired set, and its button counts what a Reset would actually bring back (shown or closed, counted once). Agent tips carry no catalog id and are untouched. Live repro (Playwright against the worktree's Vite renderer, all nine tips seeded as seen, clock fast-forwarded past settle + cooldown): origin/main re-showed "Start fresh"; fixed renderer shows nothing; a fresh user (nothing seen) still gets the first tip. |
||
|
|
1e7d29a081 |
test(gateway): exercise the shared overflow classifier instead of a copy
test_7100 replicated the phrase list inline, so it tested its own copy rather than production; point it at is_context_overflow_failure_result. Drop a dead `error=` parameter from the normalizer test helper. |
||
|
|
949724b321 |
fix(gateway): one context-overflow verdict for the reply and the transcript skip
Moving the `if response` guard below the failed branch exposed the normalizer's loose overflow predicate (bare "token"/"exceed"/"context"/"payload", or any 400 on a long session) to failed turns that carry real text: billing, rate-limit, auth and content-policy replies were rewritten to "Session too large / /compact". Hoist run_turn's stricter classifier (compression_exhausted, multi-word phrases, 400 on history > 50) into a module-level `is_context_overflow_failure_result` and use it for both the #1630 transcript skip and the user-facing rewrite, so the two can never disagree. Populated text is only rewritten when it is the bare provider envelope (`_looks_like_gateway_provider_error`) on an overflow turn; curated agent text (compression-timeout guidance, /compress hint) survives. Replaces the sanitizer-wording test with the passthrough invariants that catch the regression. |
||
|
|
613d1f7c03 | fix(gateway): keep session-too-large reply when a failed 400 still has text | ||
|
|
722bb1a674 | test(gateway): pin overflow reply when a failed 400 still has text | ||
|
|
8762a08b67 | test(desktop): verify PATH probe child kill on timeout | ||
|
|
f2b33fb4a2 | fix(desktop): kill the PATH probe child on timeout | ||
|
|
d5aaaa4a1b |
fix(tui-gateway): a hidden seed row stays out of search, a partial seed copy is rolled back, live resume counts the wire (#107562)
Two independent reviews of the seeded-create change found three more places where the newly durable hidden row, or the new create-time copy, was not handled by the same rule as the rest of the path: - Message search (dashboard search and the session_search tool) had no display_kind filter, so a hidden opening row matched a query the person never saw. The shared search predicate now skips hidden rows. - _seed_row left the fresh session row behind when the transcript copy failed after the row was committed. The first prompt's retry copies the whole seed, so a kept partial copy would be duplicated. The row is now deleted when the copy did not complete, the compensation _persist_branch applies to branch children; the first prompt then starts clean. - _live_session_payload (a resume that reuses a live session) reported message_count as the raw history length while its messages array was filtered. It now follows _resume_response: the stored size when messages are omitted, else the wire count. Tests: the two seeded-create tests now drive the first-submit path through _persist_session_row_for_submit, the function prompt.submit calls, and assert search and the reuse-live count; a third test pins the rollback (no row after a failed copy, one copy after the retry). |
||
|
|
c22a8d8e3f |
Seeded sessions survive a gateway restart and store their seed once (tui_gateway) (#107549)
* fix(tui-gateway): a seeded session is durable at create, and its seed is written once session.create accepts opening messages. Three defects sat in that path: - A seeded session without a parent was never persisted at create, so a restart before the first prompt lost it and session.resume answered 4007. Only branch children (#93959) were persisted up front. The same rationale applies to any seeded create: seeded content is intent, not an abandoned draft. Parentless seeds now persist their row, transcript and client title at create; empty drafts stay lazy. - _coerce_seed_history dropped display_kind, so a seeded row tagged "hidden" (model-facing scaffolding) rendered as a user bubble. The coercion keeps "hidden" and only "hidden"; every other kind is stamped by the gateway at turn time and is not accepted from the wire. - A branch child's seed was written twice: _seed_branch_row copied it at create but never marked it persisted, so the first prompt's _persist_branch_seed appended the copy again. The create path now sets _branch_seed_persisted, and the gate is a create-time `seeded` stamp instead of parent_session_id, so a resumed session (whose history comes from the DB) can never re-append its transcript. Two invariant tests, both red on main: a parentless seed survives a gateway restart with the hidden row kept out of the wire transcript and not re-written by the first-submit path; a branch child's seed is stored exactly once. The reasoning-fields fixture stamps `seeded`, the flag session.create sets. * fix(tui-gateway): a hidden seed row stays out of the list preview and the create count Live-testing the seeded create on every surface showed two places where the newly durable hidden row (display_kind="hidden") still surfaced: - session.list built a session's preview from its first user row with no display_kind filter, so a hidden opening row (model-facing scaffolding the gateway never paints) became the sidebar preview. The preview predicate now skips hidden rows, in every listing query that shares it. - session.create reported message_count as the raw seed length while its messages array already filtered the hidden row (2 vs 1). It now counts what is on the wire, the same rule session.resume applies. Both are covered by the existing seeded-create test: the create count equals the wire transcript, and the preview of a session whose first user row is hidden is its first visible user row. * fix(tui-gateway): a live unpersisted resume counts the wire transcript session.resume on a live session that has no row yet reported message_count as the raw history length while its messages array was already filtered, the same mismatch the previous commit fixed on session.create. Count the wire, as the cold, deferred and reuse-live resume paths already do. * chore: retrigger CI (zero-job dispatch failure, auto-heal) |
||
|
|
438a313500 | docs(honcho): document a2aSessions and per-author writes on the site page | ||
|
|
d55f1ed0a8 |
test(honcho): trim the author-peer suites to their invariants
Keep the isolation contracts (bot never on the human peer, bot turn never in the human session, writes refused mid bot-turn, one join per author, signature busts the cache) and drop the alias/prefix/sanitize enumerations. 795 -> 454 lines. |
||
|
|
a42388b3a9 |
feat(agent): a2a_key names a bot author's turns
Dropped from the #103888 salvage for lack of a consumer; the honcho a2a session key is that consumer. |
||
|
|
d74f13e4a0 |
fix(honcho): include a2aSessions in identity_signature
sync_turn reads a2a_sessions from the config bound when the provider was built. A cached gateway provider kept the old value after honcho.json flipped it, because the signature that busts that cache did not carry the flag. |
||
|
|
431cd9084b |
fix(honcho): bound the joined author peer memory by session count
_joined_author_peers kept an entry for every honcho session the manager ever wrote to. It now holds at most _SESSION_CACHE_MAX_SIZE sessions and drops the oldest past that, so a forgotten session's authors rejoin on their next write. A failed join no longer leaves an empty entry behind. |
||
|
|
bcac7e9465 |
fix(honcho): read an author join's observation flags through one manager method
The join read the manager-wide user_observe_me and user_observe_others directly. It now asks _join_observation_flags(honcho_session_id), which returns the same values today. #103889 stores the effective flags per session and replaces the body of that method. |
||
|
|
2ac7fcddf8 |
fix(honcho): a bot author never lands on the session's human runtime peer
_generated_runtime_peer_id takes a reserved set, and the bot path passes the session's human peer ids: each runtime id and the peer _resolve_user_peer_id returns for the key. bot:coder with a runtime human coder and no runtimePeerPrefix now gets the digest suffix. _explicit_user_peer_ids keeps its meaning for prefixed runtime users. |
||
|
|
a2c65ace0a |
test(honcho): cover bot:<connection>/<profile> authors end to end
Two senders named coder on different connections get different peers and different a2a sessions, and a userPeerAliases entry keyed by the full connection-qualified id wins. |
||
|
|
d96a6d9ab9 |
fix(honcho): include the workspace in identity_signature
identity_signature now carries cfg.workspace_id. The gateway folds these values into its agent cache key, and a workspace change in honcho.json reused a cached agent that was still bound to the old workspace. |
||
|
|
8704e9ca4c |
fix(honcho): put this agent's aiPeer in the a2a session key
_a2a_session_key now names the session <session>:a2a:<aiPeer>:<sender id>-<digest>. Two profiles that share a workspace and a session key wrote one sender's DMs into one Honcho session. The recipient peer comes from the same aiPeer derivation the session builder uses, moved into session_peers.assistant_peer_id_for so the two cannot drift. |
||
|
|
b4d7a33b51 |
fix(honcho): derive a bot author's peer from its full id with the runtime digest rule
_peer_id_for_runtime_id now looks up userPeerAliases by the full bot id and otherwise passes everything after bot: through _generated_runtime_peer_id. A digest suffix is added when sanitizing changed the id or the result equals peerName or an alias target, so bot:eri never resolves to the operator's peer and bot:a.b stays apart from bot:a-b. The docstring and README no longer claim a cloned profile's aiPeer defaults to the profile name. |
||
|
|
bdeb6f1b77 |
refactor(honcho): name the a2a session from core's a2a_key
The plugin spelled the `a2a:` prefix itself. `agent.turn_author.a2a_key` is the shared name for a bot author's turns, so the session key now derives from it and every reader that files bot turns apart agrees on the prefix. The resulting key is unchanged. |
||
|
|
170589dd10 |
fix(honcho): every bot author gets its own peer or its turn is skipped
A gateway platform marks a bot sender with its raw user id and a bot flag, never a `bot:` id. `resolve_author_peer_id` treated that author as a human, so `pinUserPeer` collapsed a bot onto `peerName` and an unresolved peer opened the a2a session under the human's peer. The resolver now takes `is_bot` and gives every bot its own peer. `sync_turn` skips the turn when no peer resolves or when the peer equals this agent's `aiPeer`. The a2a session key carries an eight-character digest of the author id, so two ids that sanitize alike stay in separate sessions. During a bot-authored turn `honcho_conclude` and `honcho_profile` refuse writes and the built-in memory mirror is skipped, because conclusions and cards describe the human. The README paragraph on bot DMs now matches the code. |
||
|
|
9f2a9384dc |
fix(honcho): pinUserPeer collapses the operator's accounts, not bot authors
with pinUserPeer on, resolve_author_peer_id returned None for every author, so a bot dm's words were written under the human's pinned peer inside the a2a session. the pin exists to unify one person's platform accounts. a bot is not one of them. bot: authors now resolve to their peer before the pin check, so a pinned operator still gets bot speech attributed to the bot. |
||
|
|
df1513b728 | docs(honcho): document a2aSessions and bot dm attribution | ||
|
|
f7d5ac3230 |
feat(honcho): write bot dms into their own a2a session
A DM relayed from another Hermes profile ran as a turn in the recipient's Bot Chat session. sync_turn wrote the bot's words and the recipient's reply into that session, and before per-author writes they landed under the human's peer. The human's representation absorbed conversations the human never had. The turn context now marks such turns with scope a2a:<bot id>. sync_turn routes a bot-authored turn into a separate Honcho session keyed <session>:a2a:<sanitized bot id>, created with the sender bot as its user peer, and never writes it into the human's session. The key is deterministic so every turn from the same bot reaches the same session, and it stays inside Honcho's 100 character session id limit. Recall still reads the human's session only. a2aSessions (host block, then root, default true) turns the routing on. With it off, bot-authored turns are skipped. A bot turn that names no author id is skipped as well, because nothing can key its session. Human turns are unchanged. get_or_create takes a user_peer_id override so the a2a session's roster is the bot and the assistant, not the runtime human. |
||
|
|
88fc9402d8 |
feat(honcho): declare identity_signature and drop the gateway's honcho keys
The gateway agent cache read honcho.json itself through a honcho-named block in gateway/run.py and gateway/run_agent_cache.py. Every other memory provider had no way to bust the cache when its identity mapping changed. HonchoMemoryProvider.identity_signature() now returns the same values under provider-neutral keys: user_identity, agent_identity, pin_user_identity, runtime_identity_prefix, user_identity_aliases, session_prefixing. The gateway files them under memory.<key> through the MemoryProvider hook. The hook reads config only, memoizes on the file's mtime and size, and returns an empty dict when the file cannot be read. The honcho-specific extractor, its memo and its key tuple are gone from the gateway. The pinPeerName cache-busting test now asserts on memory.pin_user_identity. |