Commit Graph

15566 Commits

Author SHA1 Message Date
Siddharth Balyan 3b01b4ce0f feat(desktop): Nous free tier on Hermes Desktop (#105260)
* feat(desktop): free-tier state over RPC, status routes that name it, and a sign-in that keeps connectors

The desktop learns about the Nous free tier by reading local auth state (pull): free_tier.status
answers has_guest / enabled / carries_inference / notice_pending with zero network, and
free_tier.ack_notice persists the one-time notice flag on the identity itself. setup.runtime_check
reports free_tier for the selected route; /api/portal, the Nous card in /api/providers/oauth and
billing.state carry free_tier (billing answers the free tier locally instead of a portal call that
can only fail). The free-tier picker row carries an explicit free_tier_row flag and is never priced or
locked. POST /api/providers/oauth/nous/start over a free-tier identity registers the connector
transfer and returns its code and consent URL; the poller waits for the transfer before the token
grant, persists the account, runs settle_after_upgrade, and the poll response gains reason,
account_email and model.

* feat(desktop): free tier on Hermes Desktop: ready screen, notice strip, status chip, Billing view, one sign-in dialog

The renderer reads the free tier from free_tier.status (pull) into one store; the first-launch
intro is the same state rendered two ways, keyed on the backend's one-time flag: the onboarding
overlay opens on a ready screen when the free tier carries inference, else a one-time strip above
the composer. Settings > Billing gains a free_tier view (notice with one Sign in, Plan / Model /
Connectors summary, plan card, footnote; no payment or usage rows). A status-bar chip names the
tier and model while it carries inference. Every entry point opens one claimed sign-in dialog that
drives the extended oauth/nous route and maps the poll's status and reason to the ruled screens;
Done settles billing, model options, providers and re-homes a session still on nous/welcome. The
picker badge also fires on free_tier_row. Docs: Desktop section in the free-tier guide, AGENTS notes.

* fix(desktop): free_tier.status starts the free tier's background setup when no identity exists

A served backend has no session-setup moment like the CLI's, so beside an explicit provider the free
tier was never set up on the desktop: no connectors, no notice strip. The first status read now
starts the same one-attempt background setup; the call itself never waits.

* fix(desktop): one Sign in on the Billing page; Settings > Providers names the free tier, never Connected

The free-tier plan card is the what-you-get text alone (the notice carries the page's one Sign in).
The Nous provider row reads Nous · free tier with a Free tier tag while the identity is the free
tier, instead of Nous Portal · Connected.

* fix(desktop): Settings > Providers never files the free tier under Connected

* fix(desktop): the intro's shape is keyed on the route, not on the identity

free_tier.status reports available (an identity exists and the tier is on); whether inference
runs on the free tier is setup.runtime_check.free_tier, keyed on the resolved endpoint. The ready
screen shows when that route is the free tier; the composer strip when the user's own provider
carries inference. An own-key install used to get the ready screen.

* docs(desktop): say what the free-tier chip is keyed on

* fix(desktop): the featured Nous row's pitch on the free tier says what signing in adds

* fix(desktop): a cancelled or superseded sign-in attempt can no longer change the identity or hide the intro

Four lifecycle holes from review. The Nous poller checks the session's cancelled flag after the
transfer wait, after the token grant, and once more under the session lock together with the
save, so a sign-in the user abandoned never persists. The renderer's sign-in store carries an
attempt generation that every continuation checks after each await, so a poll from a closed
attempt cannot publish over the one on screen (and its backend session is cancelled). The ready
screen comes down only after the backend recorded the acknowledgement. A composer still mounted
takes over the notice claim when its owner unmounts. One thin test per hole.
2026-09-11 03:45:32 +05:30
Siddharth Balyan 04a76c4109 fix(auth): a sign-in from the free tier settles the default model and route (#105259)
* fix(auth): a sign-in from the free tier settles the default model and route once, for every caller

Picking the free-tier row leaves model.default at nous/welcome pinned to the welcome host. After a
sign-in an account cannot keep either: the welcome host refuses account tokens, and the portal host
serves nous/welcome as a paid model. One completion step, anon_auth.settle_after_upgrade, now runs
after the account is persisted: a config on the free tier's route moves to the account's inference
host and the recommended default for the account's plan, through the same config write a plain
Nous login uses; a config on the user's own model is left alone. The pick is the one
GET /api/model/recommended-default already makes, factored into models.recommended_nous_default_model
so the CLI and the desktop land on the same model. hermes auth upgrade prints the new default.

* fix(auth): a sign-in completion with no eligible recommendation leaves no default model

The static provider-wide default is not narrowed by the account's plan or org policy, so writing it
as a fallback could persist a model the account may not use. When the recommendation cannot yield a
model, the route still moves to the account's host but model.default is left unset; the CLI says so
and points at `hermes model`.

* docs(free-tier): say what happens when no recommendation is available after sign-in

* fix(auth): sign-in completion moves the host and clears the default in one config write

Two writes could fail between them and leave the account host paired with nous/welcome.
_update_config_for_provider gains clear_default so the caller with no model to offer removes
model.default in the same atomic write that sets the host.
2026-09-11 03:45:31 +05:30
Siddharth Balyan a2db110ccc feat(auth): Nous free tier: free inference and connectors out of the box, one command to sign in (#105258)
* feat(auth): Nous free tier core: anonymous identity minted on first use, welcome inference, shared-store scoping

A fresh install with no provider sets up a free Nous identity (anonymous auth method of the nous
provider) instead of forcing the setup wizard. The identity is persisted through the same path a
real login uses, so the resolver ladder is unchanged. Two seams differ: token acquisition
re-exchanges the anon credential (no refresh token), and routing pins the welcome host's single
model nous/welcome. One identity per shared store; nous.guest: false turns the free tier off.

* test(auth): free tier core contracts: lifecycle, resolver precedence, exchange seam, model pin

* docs(user-guide): free tier and signing in

New page explaining what a fresh install gets before any key or sign-in
(free inference on nous/welcome plus connectors), how the free tier
coexists with a user's own API key, how to sign in with hermes auth
upgrade and keep connectors, how to turn the free tier off with
nous.guest, what hermes logout does in each state, a troubleshooting
table, and a plain privacy note. Wired into the Using Hermes sidebar.

* fix(auth): logout leaves the free tier alone and clears the shared store for a real Nous account

Logging out of the free tier is a no-op: it is not a login, so nothing is cleared and the user
is told they were never signed in. Logging out of a real Nous account now also clears the
cross-profile store, so a profile logout is not silently re-adopted on the next boot.

* fix(model): switching off the free tier points at signing in, never hops providers

* Name the free tier in the gateway startup notice and tell explicit-provider installs about it once

* Render the Nous free tier as free tier on auth status, auth list, hermes status and portal info, short-circuit billing copy for it, and skip the keepalive when there is no refresh token

* fix(auth): free tier is set up where nothing is configured: resolver last rung and first-run check

Both the provider resolver's terminal rung and the CLI first-run check now try to set up the free
tier before declaring nothing configured. On a fresh install the first command lands in chat on
nous/welcome; a failed setup still falls through to the existing guidance.

* Add hermes auth upgrade: sign the free tier into a Nous account while keeping its connectors

The device-code flow runs as usual, with a promotion intent registered on the portal between the
code request and the token poll so the account that approves the code inherits the free tier's
connectors. The promotion status decides the outcome: only a completed one is followed by the
token grant, which is persisted over the free-tier singleton and the shared store. Declined,
superseded, retired and busy outcomes each print their own plain copy, and a retired identity is
cleared so the next use sets up a fresh one. User-facing text never names the free tier's internals.

* Show the Nous free tier as one picker row with nous/welcome and hide it when nous.guest is off

* fix(auth): upgrade opens the consent page for this sign-in; one mint attempt per process; forced free tier wins the first-run check

The browser leg of hermes auth upgrade now prints and opens the promotion claim URL with the
claim code, not the generic device page. A failed mint is attempted once per process so several
bootstrap sites cannot hit a closed gate or a 429 twice; a retired credential resets that so
re-minting still happens. HERMES_FORCE_GUEST is honoured ahead of the first-run provider check.

* fix(auth): pin the welcome model on the selected route, not on profile state; background setup retries after a failure

A credential-pool entry can select a paid Nous key while the profile singleton is still the free
tier. The model pin now keys on the resolved endpoint (welcome host) in agent init and /model, and
the pin in model normalization is removed since it had no route to look at. A failed background
identity setup releases its latch so a later attempt in the same process can try again.

* fix(auth): decide the Nous model together with the route on every credential-pool swap

The credential pool can move a Nous agent between the welcome host and the portal host after
init. One helper, pin_model_for_route, now runs at init and inside every pool swap, so the
welcome host always carries nous/welcome and a paid endpoint always keeps the caller's model.

* fix(auth): apply the route model policy on every wire mode during a pool swap; release the setup latch if the thread cannot start

* fix(auth): free-tier lifecycle takes profile then shared lock, reconciles with the shared store, persists the mint before exchanging, and clears only the identity that died

The shared store is the identity of record for a Hermes root: a profile holding a stale free-tier
identity adopts a sibling's newer sign-in instead of keeping the guest, and never overwrites the
shared account. Locks are taken in the documented order (profile, then shared). A minted credential
is stored as soon as create succeeds, so a rate-limited or timed-out exchange does not lose it and
trigger a second mint. Retiring a dead credential removes only that credential from both stores.
Guest exchange uses the resolver's canonical portal URL.

* fix(auth): a credential rotation never rewrites the conversation model; connectors honour the off switch and replace a retired free-tier credential

The welcome host serves one model, so a rotation onto it is refused for any conversation on another
model instead of silently switching that conversation to nous/welcome (the model pin applies only
when a route is first chosen). The connector token path now treats the free tier as absent when
nous.guest is false, including cached tokens, and shares the one dead-credential rule with
inference: a retired identity is replaced once rather than returning its stale token.

* fix(auth): plain login never imports the free tier as OAuth credentials; the gateway startup line reads persisted state only

A free-tier identity in the shared store is not an OAuth credential to offer for import; a real
sign-in replaces it. The gateway's startup notice now answers provider precedence from persisted
state (no token refresh at boot), so an expired free-tier token cannot stall the online message.
2026-09-11 03:45:31 +05:30
SHL0MS 2ddeba9e17 Merge pull request #99523 from somewheresy/justin/e-1047-route-hermes-actual-provider-through-chat-completions-with
fix(providers): use chat completions for all Actual routes
2026-09-10 17:25:53 -04:00
Teknium 564aef2946 fix: /model onto Bedrock Mantle keeps SigV4 auth instead of 401ing
Startup (agent_init._init_openai_client) ran configure_bedrock_openai_client_kwargs,
so the aws-sdk sentinel became a SigV4-signing httpx client. Every later client
rebuild — switch_model, fallback restore, credential rotation, request-scoped
clients — went through create_openai_client with bare {api_key, base_url} kwargs,
so the OpenAI SDK sent "Authorization: Bearer aws-sdk" and Mantle answered 401
"Invalid bearer token". Symptom: `hermes --provider bedrock --model
openai.gpt-5.6-terra` works, `/model openai.gpt-5.6-terra` inside the CLI/TUI fails.

Install SigV4 in create_openai_client itself (the single chokepoint every primary
OpenAI-wire client passes through) whenever the base_url is a Mantle host, so all
rebuild paths inherit the fix rather than each remembering to call the adapter.
A real AWS_BEARER_TOKEN_BEDROCK key is left alone (the adapter only rewrites the
aws-sdk / no-key-required placeholders).

Live repro (local HTTP sink capturing the Authorization header after switch_model):
before "Bearer aws-sdk", after "AWS4-HMAC-SHA256 Credential=...".
2026-09-10 12:18:44 -07:00
Justin Bennington 8a6b5b67a7 fix(providers): block Actual at Responses send sites (E-1047) 2026-09-10 15:13:02 -04:00
Justin Bennington 4135933ec9 fix(providers): preserve Actual routing during startup and key reload (E-1047) 2026-09-10 15:01:17 -04:00
Justin Bennington 8b2b83a906 fix(providers): pin all Actual routes to chat completions (E-1047) 2026-09-10 14:56:55 -04:00
Justin Bennington d7b0a72c2a fix(providers): route Actual through chat completions (E-1047) 2026-09-10 14:56:55 -04:00
Teknium f98cb00a8b fix(vault): 2FA review follow-ups — per-digit fill only for an unmistakable maxlength=1 widget; honour otpauth digits/period/algorithm
Reviewer findings on #107585:
- build_otp_fills split any >=4 code-like controls into digits. A page with
  promo/zip/referral 'code' inputs next to the real OTP box would have had a
  digit sprayed across unrelated fields. Split now requires exactly len(code)
  controls that are all maxlength=1, same form, adjacent in DOM order
  (inspection JS exports maxLength); anything else fills ONE field, the
  best-scoring one. Verified on the real Browser Use stack: 6-box widget gets
  one digit each; scattered page fills only the one-time-code input.
- normalize_otp_secret dropped digits/period/algorithm from otpauth:// URIs,
  so an 8-digit or SHA-256 authenticator would get wrong codes. Non-default
  parameters are now stored as seed|digits|period|algo and honoured (RFC 6238
  SHA-256 8-digit vector added); hotp:// is rejected explicitly.
- rebase on main (prompts.ts conflict) + prettier.
2026-09-10 11:48:01 -07:00
Teknium d9ca9c974d feat(vault): two-factor codes — automatic from a saved authenticator key, otherwise asked for in the user's UI
Follow-up to #106480. Sites that ask for a code after the password stopped
the agent cold: the login classifier excludes one-time-code fields on
purpose (a password must never land in an OTP box) and there was no tool
for the second step, so the only move was to ask in chat.

browser_vault_enter_code
  Fills the one-time code the current page asks for. Two sources, same
  invariant as passwords (the code goes to the page over the supervisor
  socket and never enters model context):
  - a TOTP seed on the login: local vault `otp_secret` (RFC 6238, stdlib,
    verified against the RFC test vectors), 1Password `op item get --otp`,
    Bitwarden `bw get totp`. Nobody is asked.
  - no seed: the surface prompts "Verification code for {site}"; the user
    types what their phone/email/app shows. Enter on empty / Skip declines
    and the tool returns code_declined ("do not ask again this turn").
  no_code_field tells the model the site wants a passkey / hardware key /
  app approval: hand it to the user's device and wait for navigation.
  Per-digit OTP boxes (maxlength=1 pattern) get one digit each in DOM order.

Surfaces
  CLI: sudo-style panel, code shown as typed (not a secret worth masking,
  typos must be visible), Enter submits, ESC/empty skips.
  Desktop: "Verification code for {site}" card via vault.code.request /
  vault.code.respond (gateway), owner-routed like the other vault prompts.
  Settings → Passwords & Logins: optional "Authenticator key" field on the
  add form (base32 or otpauth:// link); items with one show a "2FA auto"
  badge. `hermes vault add` asks for the same optional key.
  browser_vault_fill's result now says what to do next ("if the site asks
  for a verification code, call browser_vault_enter_code with this handle").
  Six locales.

Verified live (real model, local 2FA site that checks the TOTP; CLI PTY):
  A. login saved with authenticator key → signed in through 2FA, zero
     prompts, code/password absent from the transcript
  B. login without key → code panel → user types code → signed in
  C. panel dismissed → agent stops and explains, never asks in chat
Unit: RFC 6238 vectors, seed normalisation, mint-without-asking, per-digit
spread, decline, no-code-field; Desktop card test (owner routing, trim, Skip).
2026-09-10 11:48:01 -07:00
kshitijk4poor 1e7d29a081 test(gateway): exercise the shared overflow classifier instead of a copy
test_7100 replicated the phrase list inline, so it tested its own copy rather
than production; point it at is_context_overflow_failure_result. Drop a dead
`error=` parameter from the normalizer test helper.
2026-09-10 23:45:03 +05:30
kshitijk4poor 949724b321 fix(gateway): one context-overflow verdict for the reply and the transcript skip
Moving the `if response` guard below the failed branch exposed the normalizer's
loose overflow predicate (bare "token"/"exceed"/"context"/"payload", or any 400
on a long session) to failed turns that carry real text: billing, rate-limit,
auth and content-policy replies were rewritten to "Session too large / /compact".

Hoist run_turn's stricter classifier (compression_exhausted, multi-word phrases,
400 on history > 50) into a module-level `is_context_overflow_failure_result` and
use it for both the #1630 transcript skip and the user-facing rewrite, so the two
can never disagree. Populated text is only rewritten when it is the bare provider
envelope (`_looks_like_gateway_provider_error`) on an overflow turn; curated agent
text (compression-timeout guidance, /compress hint) survives. Replaces the
sanitizer-wording test with the passthrough invariants that catch the regression.
2026-09-10 23:45:03 +05:30
xxxigm 722bb1a674 test(gateway): pin overflow reply when a failed 400 still has text 2026-09-10 23:45:03 +05:30
Siddharth Balyan d5aaaa4a1b fix(tui-gateway): a hidden seed row stays out of search, a partial seed copy is rolled back, live resume counts the wire (#107562)
Two independent reviews of the seeded-create change found three more
places where the newly durable hidden row, or the new create-time copy,
was not handled by the same rule as the rest of the path:

- Message search (dashboard search and the session_search tool) had no
  display_kind filter, so a hidden opening row matched a query the
  person never saw. The shared search predicate now skips hidden rows.
- _seed_row left the fresh session row behind when the transcript copy
  failed after the row was committed. The first prompt's retry copies
  the whole seed, so a kept partial copy would be duplicated. The row
  is now deleted when the copy did not complete, the compensation
  _persist_branch applies to branch children; the first prompt then
  starts clean.
- _live_session_payload (a resume that reuses a live session) reported
  message_count as the raw history length while its messages array was
  filtered. It now follows _resume_response: the stored size when
  messages are omitted, else the wire count.

Tests: the two seeded-create tests now drive the first-submit path
through _persist_session_row_for_submit, the function prompt.submit
calls, and assert search and the reuse-live count; a third test pins
the rollback (no row after a failed copy, one copy after the retry).
2026-09-10 18:05:33 +00:00
Siddharth Balyan c22a8d8e3f Seeded sessions survive a gateway restart and store their seed once (tui_gateway) (#107549)
* fix(tui-gateway): a seeded session is durable at create, and its seed is written once

session.create accepts opening messages. Three defects sat in that path:

- A seeded session without a parent was never persisted at create, so a
  restart before the first prompt lost it and session.resume answered
  4007. Only branch children (#93959) were persisted up front. The
  same rationale applies to any seeded create: seeded content is
  intent, not an abandoned draft. Parentless seeds now persist their
  row, transcript and client title at create; empty drafts stay lazy.
- _coerce_seed_history dropped display_kind, so a seeded row tagged
  "hidden" (model-facing scaffolding) rendered as a user bubble. The
  coercion keeps "hidden" and only "hidden"; every other kind is
  stamped by the gateway at turn time and is not accepted from the wire.
- A branch child's seed was written twice: _seed_branch_row copied it at
  create but never marked it persisted, so the first prompt's
  _persist_branch_seed appended the copy again. The create path now
  sets _branch_seed_persisted, and the gate is a create-time `seeded`
  stamp instead of parent_session_id, so a resumed session (whose
  history comes from the DB) can never re-append its transcript.

Two invariant tests, both red on main: a parentless seed survives a
gateway restart with the hidden row kept out of the wire transcript and
not re-written by the first-submit path; a branch child's seed is stored
exactly once. The reasoning-fields fixture stamps `seeded`, the flag
session.create sets.

* fix(tui-gateway): a hidden seed row stays out of the list preview and the create count

Live-testing the seeded create on every surface showed two places where
the newly durable hidden row (display_kind="hidden") still surfaced:

- session.list built a session's preview from its first user row with no
  display_kind filter, so a hidden opening row (model-facing scaffolding
  the gateway never paints) became the sidebar preview. The preview
  predicate now skips hidden rows, in every listing query that shares it.
- session.create reported message_count as the raw seed length while its
  messages array already filtered the hidden row (2 vs 1). It now counts
  what is on the wire, the same rule session.resume applies.

Both are covered by the existing seeded-create test: the create count
equals the wire transcript, and the preview of a session whose first
user row is hidden is its first visible user row.

* fix(tui-gateway): a live unpersisted resume counts the wire transcript

session.resume on a live session that has no row yet reported message_count as
the raw history length while its messages array was already filtered, the same
mismatch the previous commit fixed on session.create. Count the wire, as the
cold, deferred and reuse-live resume paths already do.

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-10 17:55:24 +00:00
Teknium d55f1ed0a8 test(honcho): trim the author-peer suites to their invariants
Keep the isolation contracts (bot never on the human peer, bot turn never in the
human session, writes refused mid bot-turn, one join per author, signature busts
the cache) and drop the alias/prefix/sanitize enumerations. 795 -> 454 lines.
2026-09-10 10:45:57 -07:00
Erosika d74f13e4a0 fix(honcho): include a2aSessions in identity_signature
sync_turn reads a2a_sessions from the config bound when the provider was built. A cached gateway provider kept the old value after honcho.json flipped it, because the signature that busts that cache did not carry the flag.
2026-09-10 10:45:57 -07:00
Erosika 431cd9084b fix(honcho): bound the joined author peer memory by session count
_joined_author_peers kept an entry for every honcho session the manager ever wrote to. It now holds at most _SESSION_CACHE_MAX_SIZE sessions and drops the oldest past that, so a forgotten session's authors rejoin on their next write. A failed join no longer leaves an empty entry behind.
2026-09-10 10:45:57 -07:00
Erosika bcac7e9465 fix(honcho): read an author join's observation flags through one manager method
The join read the manager-wide user_observe_me and user_observe_others directly. It now asks _join_observation_flags(honcho_session_id), which returns the same values today. #103889 stores the effective flags per session and replaces the body of that method.
2026-09-10 10:45:57 -07:00
Erosika 2ac7fcddf8 fix(honcho): a bot author never lands on the session's human runtime peer
_generated_runtime_peer_id takes a reserved set, and the bot path passes the session's human peer ids: each runtime id and the peer _resolve_user_peer_id returns for the key. bot:coder with a runtime human coder and no runtimePeerPrefix now gets the digest suffix. _explicit_user_peer_ids keeps its meaning for prefixed runtime users.
2026-09-10 10:45:57 -07:00
Erosika a2c65ace0a test(honcho): cover bot:<connection>/<profile> authors end to end
Two senders named coder on different connections get different peers and different a2a sessions, and a userPeerAliases entry keyed by the full connection-qualified id wins.
2026-09-10 10:45:57 -07:00
Erosika d96a6d9ab9 fix(honcho): include the workspace in identity_signature
identity_signature now carries cfg.workspace_id. The gateway folds these values into its agent cache key, and a workspace change in honcho.json reused a cached agent that was still bound to the old workspace.
2026-09-10 10:45:57 -07:00
Erosika 8704e9ca4c fix(honcho): put this agent's aiPeer in the a2a session key
_a2a_session_key now names the session <session>:a2a:<aiPeer>:<sender id>-<digest>. Two profiles that share a workspace and a session key wrote one sender's DMs into one Honcho session. The recipient peer comes from the same aiPeer derivation the session builder uses, moved into session_peers.assistant_peer_id_for so the two cannot drift.
2026-09-10 10:45:57 -07:00
Erosika b4d7a33b51 fix(honcho): derive a bot author's peer from its full id with the runtime digest rule
_peer_id_for_runtime_id now looks up userPeerAliases by the full bot id and otherwise passes everything after bot: through _generated_runtime_peer_id. A digest suffix is added when sanitizing changed the id or the result equals peerName or an alias target, so bot:eri never resolves to the operator's peer and bot:a.b stays apart from bot:a-b. The docstring and README no longer claim a cloned profile's aiPeer defaults to the profile name.
2026-09-10 10:45:57 -07:00
Erosika bdeb6f1b77 refactor(honcho): name the a2a session from core's a2a_key
The plugin spelled the `a2a:` prefix itself. `agent.turn_author.a2a_key` is the shared name
for a bot author's turns, so the session key now derives from it and every reader that files
bot turns apart agrees on the prefix. The resulting key is unchanged.
2026-09-10 10:45:57 -07:00
Erosika 170589dd10 fix(honcho): every bot author gets its own peer or its turn is skipped
A gateway platform marks a bot sender with its raw user id and a bot flag, never a `bot:` id.
`resolve_author_peer_id` treated that author as a human, so `pinUserPeer` collapsed a bot onto
`peerName` and an unresolved peer opened the a2a session under the human's peer. The resolver now
takes `is_bot` and gives every bot its own peer. `sync_turn` skips the turn when no peer resolves
or when the peer equals this agent's `aiPeer`.

The a2a session key carries an eight-character digest of the author id, so two ids that sanitize
alike stay in separate sessions. During a bot-authored turn `honcho_conclude` and `honcho_profile`
refuse writes and the built-in memory mirror is skipped, because conclusions and cards describe
the human. The README paragraph on bot DMs now matches the code.
2026-09-10 10:45:57 -07:00
Erosika 9f2a9384dc fix(honcho): pinUserPeer collapses the operator's accounts, not bot authors
with pinUserPeer on, resolve_author_peer_id returned None for every author,
so a bot dm's words were written under the human's pinned peer inside the
a2a session. the pin exists to unify one person's platform accounts. a bot
is not one of them.

bot: authors now resolve to their peer before the pin check, so a pinned
operator still gets bot speech attributed to the bot.
2026-09-10 10:45:57 -07:00
Erosika f7d5ac3230 feat(honcho): write bot dms into their own a2a session
A DM relayed from another Hermes profile ran as a turn in the recipient's
Bot Chat session. sync_turn wrote the bot's words and the recipient's reply
into that session, and before per-author writes they landed under the
human's peer. The human's representation absorbed conversations the human
never had.

The turn context now marks such turns with scope a2a:<bot id>.
sync_turn routes a bot-authored turn into a separate Honcho session keyed
<session>:a2a:<sanitized bot id>, created with the sender bot as its user
peer, and never writes it into the human's session. The key is deterministic
so every turn from the same bot reaches the same session, and it stays
inside Honcho's 100 character session id limit. Recall still reads the
human's session only.

a2aSessions (host block, then root, default true) turns the routing on.
With it off, bot-authored turns are skipped. A bot turn that names no
author id is skipped as well, because nothing can key its session. Human
turns are unchanged.

get_or_create takes a user_peer_id override so the a2a session's roster is
the bot and the assistant, not the runtime human.
2026-09-10 10:45:57 -07:00
Erosika 88fc9402d8 feat(honcho): declare identity_signature and drop the gateway's honcho keys
The gateway agent cache read honcho.json itself through a honcho-named
block in gateway/run.py and gateway/run_agent_cache.py. Every other memory
provider had no way to bust the cache when its identity mapping changed.

HonchoMemoryProvider.identity_signature() now returns the same values under
provider-neutral keys: user_identity, agent_identity, pin_user_identity,
runtime_identity_prefix, user_identity_aliases, session_prefixing. The
gateway files them under memory.<key> through the MemoryProvider hook. The
hook reads config only, memoizes on the file's mtime and size, and returns
an empty dict when the file cannot be read.

The honcho-specific extractor, its memo and its key tuple are gone from the
gateway. The pinPeerName cache-busting test now asserts on
memory.pin_user_identity.
2026-09-10 10:45:57 -07:00
Erosika 46d625b097 feat(honcho): map bot authors onto their profile peer
The bot-mode dispatcher names another profile as bot:<profile>. The
resolver treated that like a human runtime id, so a configured
runtimePeerPrefix produced peers like telegram_bot:coder and a profile
that already owns an AI peer in the same workspace got a second one.

A bot:<profile> author now resolves in this order: a userPeerAliases entry
for the full bot id, else the sanitized profile name. A cloned profile's
aiPeer defaults to the profile name, so a same-workspace sender lands on
its existing AI peer. Prefixes never apply to bot ids. pinUserPeer still
collapses bot authors onto the pinned peer, the same as every other author.
2026-09-10 10:45:57 -07:00
Erosika 20b117ad3b fix(honcho): read the turn author from sync_turn and treat the alt id as the session peer
sync_turn only knew the author through the on_turn_start stash. The memory
manager now passes turn_author and scope with the turn, and a caller
that skips on_turn_start left the stash empty or stale.

sync_turn takes both keywords and reads the author from turn_author first.
The stash stays as the fallback for callers that never pass it.

resolve_author_peer_id compared the author against the primary runtime id
only. A transport that names the participant by the alt id (Telegram
username instead of UID) got a second peer for the same person. Either
runtime id now counts as the session's own participant.
2026-09-10 10:45:57 -07:00
Erosika 6aff2fc65b feat(honcho): write each turn under its author's peer
The manager resolved one user peer in `get_or_create` and froze it onto the
session, then `_flush_session` chose between it and the assistant peer by
role. Every user turn in a shared session landed on that one peer, so the
first person to message the agent collected everyone else's facts — and a
Honcho conclusion, once derived, is not self-correcting.

`resolve_author_peer_id` maps the turn's author onto its own peer using the
alias-then-prefix order `_resolve_user_peer_id` already applies, so an
aliased account reaches the same peer whichever turn it wrote. `sync_turn`
resolves it before starting the write thread, so a following turn cannot
retag a queued write. `_flush_session` then writes each user message under
that peer.

A shared session's roster is open — people and other agents arrive after
the session exists — so `_author_peer_for_session` joins a peer when it
first writes instead of enumerating participants at init. Joins are
remembered per session, and a failed join still writes under the right
peer, losing only the observe config.

Three cases return None and keep the session's own peer: no author named,
the author IS the session's peer, and `pinPeerName` set — that flag is an
explicit request to unify identities, so it still collapses authors in a
shared chat.

An unnamed author stays unattributed rather than defaulting to the owner.
That preserves today's behavior for the transports that send no author, so
those turns still reach the session peer; #83500 owner-gated the memory-file
migration for the same reason.

Display names never become peer IDs — they are attacker-influenceable on
any platform where participants set their own name.
2026-09-10 10:45:57 -07:00
Teknium 96fbc47f14 fix(vault): dogfood fixes — offer save-login on every backend, keep the model off passwords, bind to the login tab
Found by using the feature as a user (natural prompts, real sites, CLI PTY + native Electron), not by naming tools:

- browser_vault_save_login was registered but never offered: toolsets.py is a hand-maintained list. Added, with an
  invariant test that every registered browser_vault_* tool is in the browser toolset.
- Vault tools were absent on the DEFAULT backend (Browser Use): the gate deferred to check_browser_requirements(),
  which is False by design there. Gate = is_browser_use_cli_mode() or check_browser_requirements().
- The model typed a page-shown demo password with browser_type and offered to take one in chat: the vault rules
  lived only on the vault tools. browser_type/browser_exec now carry a vault note when the vault tools are
  present ("call browser_vault_list first … never type a password with this tool, never accept one in chat, even
  if the page shows it"); the browser_exec login-wall line points at the vault instead of "ask the user".
- On Browser Use the saved item was bound to chrome://new-tab-page: the supervisor's default page session is the
  daemon's blank tab. browser_vault_save_login now focuses the tab holding a password field before reading its
  origin (focus_page("", accept=probe); about:/chrome: pages are never candidates). Live E2E leg added.
- Settings row: "identifier · Added <date>", origin omitted when it duplicates the label.

Live (real model): CLI on Browser Use — first visit prompts, signs in, saves; second visit fills silently; GitHub
decline (Enter or ESC) stops the agent, which refuses chat passwords. CLI on the built-in stack — same three
scenarios pass. Desktop native Electron — same three scenarios plus Settings list/remove pass. Password never in
a transcript, UI, or a file outside vault/.
2026-09-10 10:35:07 -07:00
Teknium 98d11c95f4 feat(vault): zero-setup UX — save a login on the page that needs it, managers auto-detected, one "Passwords & Logins" surface
Nobody should have to learn `hermes vault add` or find a toggle before "log into GitHub" works.

- browser_vault_save_login: when the agent reaches a sign-in page with no saved login it asks the user
  on THEIR surface (CLI two-step panel on the sudo modal: identifier shown, password masked; Desktop
  card with labelled Email/username + Password fields). The answer goes to the encrypted vault bound to
  the page origin and is filled at once; the model gets back only the handle and identifier. Declining
  returns save_declined; headless sessions get prompt_unavailable. Never a password in chat.
- Vault tools ride with the browser toolset (check_browser_requirements) instead of appearing only once
  the vault has items — an empty vault is exactly when save_login is needed. browser_vault_list hints
  at it when empty.
- 1Password / Bitwarden are login sources as soon as their CLI is installed; `vault.<name>.enabled`
  is opt-OUT only. Settings shows Detected/Locked/Unlocked/Off/Not detected with a switch only for
  installed managers; `hermes vault sources` reports detection, `--disable`/`--enable` flip the opt-out.
- Desktop nav/page renamed "Passwords & Logins"; empty state tells the user they do not need to add
  anything; all five locales updated. Docs rewritten from "how it works" to "say log into X".
- New per-thread SaveLoginPrompt callback (agent/vault_backends/unlock.py) installed beside the unlock
  prompt on every CLI site and the gateway bridge (vault.save_login.request/respond/expire), propagated
  to worker threads via tools.thread_context.

Live: CLI PTY (real model, packaged Chromium, local login server) — panel shown, identifier + masked
password typed, server received the correct password, password absent from terminal transcript and
from every file under HERMES_HOME outside vault/. Native Electron (headless, isolated HOME/HERMES_HOME,
own Vite + CDP port) — card shown, "Save & sign in", server received the password, Settings lists the
saved item, password absent from the rendered UI.
2026-09-10 10:35:07 -07:00
Teknium f5e28f8119 fix(vault): declare tool parameters the registry understands; attach a supervisor to local built-in sessions
Found by a model-driven live run (hermes chat -q against a real login page): the agent found the
vault tools, typed the identifier, then failed twice for reasons the direct-call E2E could not see.

- The three schemas spelled their arguments `input_schema` (Anthropic shape). The registry and every
  provider adapter read `parameters`, so the model was shown browser_vault_fill with NO arguments and
  called it with an empty handle. Renamed; a registry-wide invariant test now fails on any schema
  without `parameters`.
- A local built-in session (agent-browser --session) carries no cdp_url, so nothing ever started a
  supervisor for it and the fill refused with supervisor_required. _ensure_supervisor asks the daemon
  for the packaged Chromium's endpoint (`get cdp-url`, carries no secret) and attaches on demand; the
  fail-closed test now pins that the daemon is only ever asked `get`, never handed an eval with the
  password.

Live: same run after the fix -> browser_navigate, browser_vault_list, browser_type, browser_vault_fill,
browser_click; the test server received the correct password; landing title "Welcome"; the password
string is absent from the whole transcript.
2026-09-10 10:35:07 -07:00
Teknium c8998c3957 feat(vault): fill payment cards and addresses at checkout, cards behind a confirm prompt
payment and address items could be stored (CLI wizard, Desktop dialog) but nothing could
fill them: a dead surface holding real card numbers. browser_vault_fill now handles all
three kinds through the same origin-bound, supervisor-only, redacted path:

- classify_checkout_control / select_checkout_fills map WHATWG autocomplete tokens
  (cc-number, cc-exp[-month|-year], cc-csc, address-line1/2, address-level1/2, postal-code,
  country-name) with label/name heuristics as backup; a combined "MM/YY" control gets
  exp_month+exp_year and suppresses the split fills; inspection now covers <select>
  (country, state, expiry month) and the fill script picks an option by value or text.
- Every payment fill goes through request_elicitation_consent (gateway button round-trip
  or CLI panel) before a byte is written; declined → payment_declined, headless sessions
  are refused. A prompt injection that reaches a checkout can ask, not spend. Card values
  join the redaction registry like passwords; the result lists targeted field tokens only.
- Origin is now required for every kind (CLI wizard asks; Desktop dialog always shows the
  field) because a card without a bound origin is unfillable.
- The tool descriptions, docs and CLI copy drop "Phase 1 / login only".

Live (evals/vault_fill_live_e2e.py, real browser_exec + packaged Chromium): decline writes
nothing; accept fills card/expiry/CVC on the /checkout tab, leaves the email box and the
country <select> untouched, and neither the card number nor the CVC appears in any result.

browser_vault_tool also: focuses the tab on the bound origin holding the right form before the
origin pre-check (focus_page from the previous commit); tool descriptions say "the browser's
input tool" (rewritten per session by model_tools); _check_vault_available is registered
uncached because its answer is per profile (vault dir + config) and the probe is a file stat.
2026-09-10 10:35:07 -07:00
Teknium 1bd4e33d36 fix(vault): make browser_vault_fill work on the default Browser Use backend
On the default backend (browser.backend unset → browser_exec) the vault tools were
advertised but could never fill: the CDP supervisor that carries the secret-bearing
eval is started only by the built-in browser_* session path, so _eval_js_secret
failed closed with supervisor_required and the origin pre-check fell back to an
agent-browser CLI eval against a browser browser_exec never touched.

- browser_exec now attaches SUPERVISOR_REGISTRY to the CDP endpoint it just routed
  the harness to (BU_CDP_WS/BU_CDP_URL), so the fill talks to the SAME browser over
  the same secret-capable WebSocket. BU direct-cloud (BU_AUTOSPAWN) exposes no
  endpoint and keeps the supervisor_required refusal.
- CDPSupervisor.focus_page(origin, accept=<js>) (used by browser_vault_fill, next commits) re-attaches the page session to the
  open tab on the item's origin whose DOM holds the form being filled (browser_exec
  opens its own tabs; the supervisor's initial attach picks the first page target,
  which is chrome://new-tab-page). browser_vault_fill uses it before the origin
  pre-check with a per-kind probe (password input / card fields / address fields).

Live: evals/vault_fill_live_e2e.py drives the real browser_exec tool against Hermes'
packaged Chromium with the login page in the third tab; A/B with the attach line
disabled fails at "did not attach a supervisor", enabled fills the password into the
/login tab and card fields into the /checkout tab with every model-facing read
scrubbed.

Also: browser_vault_list/fill described the workflow as "type the identifier with
fill_input", a helper that exists only inside browser_exec code (toolset browser-use) and is
a ghost on the built-in stack. model_tools._rewrite_browser_vault substitutes the concrete
name from the session's actual tool set (`fill_input` inside browser_exec, or browser_type),
the same dynamic cross-reference pattern browser_navigate uses for web_search.
2026-09-10 10:35:07 -07:00
Teknium 4140901d15 fix(browser): stop refusing credential-named query params on cloud browser/extract backends
browser_navigate / browser_exec / web_extract refused any URL whose query carried a
credential-NAMED parameter (token, signature, access_token, ...) when the backend was
a cloud provider. That is exactly the shape of magic links, OAuth callbacks and signed
CDN assets, so on Browserbase/Browser Use the agent could not finish a sign-in flow or
open an X video asset ("Blocked: URL contains a credential-like query parameter").

The floor protected nothing: the cloud browser already sees every cookie and typed
password of the session, and with the credential vault it receives the real password at
fill time. Hermes' own secrets leaking into a URL stay blocked by the value-shaped
_PREFIX_RE check (_secret_url_error), which is backend-independent. IMDS and
private-address floors are unchanged.
2026-09-10 10:35:07 -07:00
Teknium dedc99ec6c fix(vault): ownership and race findings from the second independent review
Manager tokens: lock generation fence (a Lock acknowledged while `bw unlock`
/ `op signin` is still running discards the late token); tokens record the
unlocking gateway session and are released when THAT session ends, not when
any sibling session in the profile is torn down.

1Password: OP_CONNECT_HOST/TOKEN come from the profile's scoped secret store
like the service token (Connect outranks a service token inside op), never
from the launch environment.

Vault RPCs bind params.profile (home + secret scope) so a shared remote
backend serving several profiles locks/lists/unlocks the requested one;
unknown profile → RPC error, not a crash.

Fill target: inspection stamps are `<nonce>:<index>`; a fill resolves only
its own inspection's stamps, so an interleaved second inspection can no
longer redirect A's password into a newly mounted field (real Chrome: 0
filled, both fields empty).

Desktop Settings: every RPC goes through the owner profile's socket
(requestGatewayForProfile), query keys carry (connection, profile), an owner
change closes dialogs and wipes drafts (a master password typed for A is
never submitted to B; a late list from A never paints under B), and vault.add
secrets travel in a ref consumed by the mutationFn instead of mutation
variables. Three owner-routing invariant tests on the real component.

Docs/PR body: session-scoped release, lock-race semantics, bw --passwordenv.
2026-09-10 10:35:07 -07:00
Teknium 5aa9e7f033 fix(vault): independent-review findings — vendor contracts, profile scope, transport, target binding
Bitwarden unlock now uses the CLI's documented non-interactive channel:
`bw unlock --raw --nointeraction --passwordenv VAR`, VAR set on the child
environment only (bw 2026.x rejects a piped password with "Master password
is required"). Verified against the real published binary.

Manager session tokens are keyed by (profile home, backend): a Desktop
gateway hosting several profiles can no longer reuse or lock another
profile's session. Status probes (`vault.sources`, is_unlocked) no longer
refresh the idle TTL; only real manager calls do. Gateway session teardown
locks the profile's managers (a per-session unlock ends with the session).

1Password service-account token comes from the profile-scoped secret store
(get_secret), not ambient os.environ.

`vault.source.set` no longer references a module constant (bind_module
rebinding dropped it → NameError on every Settings toggle).

Fill target binding: inspection stamps each input with a per-inspection
slot attribute; the fill resolves by stamp and requires type=password, then
strips every stamp. A DOM reflow between inspect and fill can no longer
redirect the password into a text field (reproduced in real Chrome before,
0 filled after).

Redaction boundary: no 4-char floor, CR/LF-normalized form registered
(what a text input actually stores), JSON object KEYS scrubbed in both
browser redactors; longest value first. Docs now state the real trust
model: accidental-disclosure protection, not an execution sandbox.

Desktop: the mid-turn card sends the master password through the owning
session's socket (requestForOwnedSession), never the ambient foreground
gateway; `vault.unlock.expire` clears a stale card; Settings keeps the
master password out of react-query mutation variables (ref consumed by the
mutationFn). One renderer invariant test for the routing.
2026-09-10 10:35:07 -07:00
Teknium b51da65258 fix: adapt execute_code cell authority to the widened prompt-callback table
_callback_api() now yields (getter, setter) pairs for every per-thread prompt
(approval, sudo, vault unlock); the kernel cell captured and restored the old
fixed 4-tuple. Iterate the table so a cell carries every callback and a future
addition needs no change here. Test recorder unpacks the new shape.

Also: perfectionist import order in ui-tui interfaces.ts (CI lint).
2026-09-10 10:35:07 -07:00
Teknium 8e7e9be217 feat(vault): sign in with 1Password or Bitwarden logins, unlocked per session
The browser vault now draws from three login sources behind one handle
shape: the local encrypted vault (vault_…), 1Password Login items (op:…)
and Bitwarden Password Manager logins (bw:…). browser_vault_list aggregates
metadata across them; browser_vault_fill routes by prefix and resolves the
password at fill time only, through the manager CLI.

External managers are locked until the user unlocks them for the current
session. The new browser_vault_unlock tool (and the fill path, implicitly)
asks the surface to show a masked master-password prompt — CLI panel
(reuses the sudo panel state), TUI/Desktop via a vault.unlock.request
blocking card. The password goes to `op signin --raw` / `bw unlock --raw`
on stdin, never argv or env; only the session token is kept, in memory,
with a 30-minute idle TTL, cleared on session close or `vault.lock`.

Headless contexts (cron, webhook, api_server, -q) can never prompt: the
manager is reported as locked with unlock=unavailable_in_this_session and
fill refuses — the same posture approvals take where nobody can answer.

Config: vault.onepassword / vault.bitwarden {enabled, binary_path, …};
a 1Password service-account token skips the prompt for headless use.
RPC: vault.sources, vault.source.set, vault.unlock, vault.lock for Settings.

Tests (2, real subprocess against a fake bw; each proven red by sabotage):
headless never prompts or spawns; unlock feeds stdin only, token never
enters os.environ, fill routes by prefix and the password only reaches the
fill script.
2026-09-10 10:35:07 -07:00
Teknium cde0ad0edd feat: agent signs into sites from an encrypted local vault (CLI, browser fill, Desktop Settings)
Consolidated re-apply of #96988 onto current main. Ported from
Merit-Systems/OpenInstinct (MIT) opaque-handle autofill design: the model
sees vault handles + login metadata, the password is resolved and filled
server-side over the supervised CDP socket, and filled values are scrubbed
from every browser tool result by an unconditional redaction registry.

Rebase adaptations to the Sep-2026 facade/sibling layout:
- toolsets: one _HERMES_CORE_TOOLS entry (the browser toolset derives from it)
- hermes_cli/main.py: vault parser registered via the subcommand owner table
- file_safety: vault/ joins the _READ_DENIED_DIRS credential-dir table
- redact: registry scrub runs before the redact_secrets early-return
- browser_vault_tool: _run_browser_command now lives in browser_tool_session
2026-09-10 10:35:07 -07:00
Teknium 94f77dfa0d fix(agent): thread turn_author through conversation_loop.run_conversation
The facade forwarded turn_author= to the loop's public entry point, which did not
declare it: every real AIAgent.run_conversation() turn raised TypeError (16 CI
failures across provider, sidecar, cron and finite-chat suites). The PR's tests
only exercised build_turn_context directly, so the missing hop was invisible.
Adds one facade-through-loop test that goes red when the kwarg is dropped.
2026-09-10 10:27:07 -07:00
Teknium 103fe14cb4 docs(gateway): document gateway.bot_loop_guard where ALLOW_BOTS is explained
The Discord page said there was no circuit breaker for bot ack-loops; there is one now.
Also strips two trailing blank lines left by the test trim.
2026-09-10 10:27:07 -07:00
Teknium ff90eb28ff test(bot-mode): trim the author and loop-guard suites to their invariants
Keep one or two behaviour tests per seam (author reset on a cached agent, forged
_turn_author refused, guard trips and cools, single charge on the busy path) and
drop the parser/setting enumerations. a2a_key goes with them: nothing in this PR
reads it; the honcho follow-up that does can bring it back with its consumer.
2026-09-10 10:27:07 -07:00
Erosika 4f12985cd2 fix(gateway): refuse relay sender fields from a logged-in client and say what the author trusts
bot_relay.deliver accepted from_profile, from_handle and from_connection from any admitted JSON-RPC client. The handler now refuses them with error 4095 when the calling transport carries a browser login identity, since a logged-in browser never relays for another connection. The DeliveryAuthor docstring now says the author is trusted because an admitted client relays it, not because the sender is verified.
2026-09-10 10:27:07 -07:00
Erosika 091ac34ba9 fix(bot-mode): qualify the peer-dm author id with the sender's hostname
message_agent sent a bare bot:<profile> id through hermes peer dm, so a remote coder and the recipient's own coder shared one author id. The peer branch now sends bot:<hostname>/<profile>, with the hostname cleaned like any author field and slashes dropped. The direct local path keeps the bare id.
2026-09-10 10:27:07 -07:00
Erosika 5ad6225174 test(bot-mode): a human prompt after a relayed dm carries no author
The author travels on the queued entry and the _run_prompt_submit call, and turn_context resets agent._turn_author at every turn start. The test pins that the human prompt following a drained relayed dm reaches run_conversation without a turn_author.
2026-09-10 10:27:07 -07:00