Commit Graph

35442 Commits

Author SHA1 Message Date
teknium1 a4f8f91e50 ci: run the OSV lockfile scan weekly against main, not on every PR
The scan is detection-only and its findings are the repo-wide baseline of
CVEs in pinned dependencies — identical for every PR, unrelated to any
PR's diff. Reporting that baseline in each PR's review comment read as
"this PR has 76 vulnerabilities" to contributors, and the SARIF upload
tripped GitHub's per-installation API rate limit during merge trains.

The scheduled weekly run (plus workflow_dispatch) keeps feeding the
Security tab; the per-PR workflow_call, the review_status wrapper job,
and the orchestrator's now-unneeded SARIF permissions are removed.
2026-09-15 11:43:45 -07:00
teknium1 f13a87e610 ci: advisory profile-scope pattern lint on the lines a PR adds
scripts/check_profile_scope_patterns.py runs the validated hazard regexes in
scripts/ci/profile_scope_patterns.json (18 of the 31 campaign patterns: every one has a scope_hint
and hits <= 50 sites on main; the wider ones are review greps, not lint) against the lines added
vs the PR base and prints file:line, pattern id/class and why. Always exits 0: most shapes have
legitimate sites (a standalone `hermes -p x` process where environ IS the profile), so the
reviewer reads each finding against its scope hint. Wired into lint.yml beside the public-surface
diff with continue-on-error.

Proof: the pre-fix tools/bot_relay.py (`env = dict(os.environ)`, before the served_profile_child_env
change) is flagged as P05/C2; the fixed file and this branch's diff vs main report 0 findings.
Test: a fixture with the hazard is flagged on the right lines, the scoped-builder version is not,
and the line filter hides hits outside the added range.
2026-09-15 10:59:22 -07:00
teknium1 3272fb35aa docs: profile-scope invariant in AGENTS.md — one process serves many profiles; out-of-turn code binds its scope
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.

Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).

Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
2026-09-15 10:59:22 -07:00
teknium1 804707bea6 fix: checkpoint store gc never runs inside a tool call or gateway startup
Symptom: `hermes update` sat for ~40s after "Refreshing cua-driver" and ended
with "Fleet version check returned no rows" (exit 1); the restarted gateway
took 26s to reach "Starting Hermes Gateway" instead of the usual 3s. The
gateway constructor was running `maybe_auto_prune_checkpoints` synchronously,
before the control socket, adapters and the code_sha stamp, and on a 1.2 GB
store its `git gc --prune=now` (a full repack) takes 20-28s — twice, because
the size-cap shrink gc'd again even when it could drop nothing.

The same defect sat on the tool-call path: `CheckpointManager._take` ran
`_enforce_size_cap`, whose `_shrink_store_to_cap` returned True without
dropping anything and triggered a 20-28s gc on the first file-mutating tool
call of every turn once the store was over the cap. That loop also re-measured
a pack size that cannot move without a gc, so a single over-cap checkpoint
dropped 20 rounds of history and flattened every project to one snapshot.

- `_take` never gcs: `_prune` and `_enforce_size_cap` rewrite refs (cheap),
  drop at most one snapshot round, and mark the store `.gc-pending`.
- `prune_checkpoints` gcs only when a ref moved (project deleted, or the
  pending marker), and its cap loop is drop -> gc -> re-measure.
- `maybe_auto_prune_checkpoints` claims the interval marker before the run
  so a failing prune costs one day, not a gc per housekeeping tick.
- `auto_prune_from_config` is the one config-driven entry point; the gateway
  calls it from the housekeeping tick (last chore), the CLI from a daemon
  thread. Nothing on either startup path waits for git.

Live A/B on a copy of a real 1.2 GB / 224-ref store: checkpoint 20.5s ->
1.2-1.6s (0 inline gc); the single repack (19.6s) now runs in the prune.
2026-09-15 10:57:16 -07:00
brooklyn! 15650b7636 test(state): cover snapshot isolation in both journal modes 2026-09-15 12:33:45 -05:00
brooklyn! 673d33095b fix(desktop): keep code geometry stable while highlighting loads 2026-09-15 12:33:45 -05:00
brooklyn! 3ccad12c0a fix(desktop): preserve reading intent across history and pane changes
Co-authored-by: wukangcheng1994 <160389295+wukangcheng1994@users.noreply.github.com>
Co-authored-by: Zhuzewen <zhuzw@seer-robotics.ai>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: nftpoetrist <264138787+nftpoetrist@users.noreply.github.com>
Co-authored-by: aydnOktay <xaydinoktay@gmail.com>
Co-authored-by: FalconOrtiz <falcon.ortiz11@gmail.com>
Co-authored-by: Ahmett101 <ahmet.tunc@gmail.com>
Co-authored-by: Val Alexander <val@opencoven.ai>
Co-authored-by: networthexplained <networthexplained@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-15 12:33:45 -05:00
brooklyn! 76f9bf948b fix(desktop): retain scoped transcript paint through background hydration
Co-authored-by: Benjamin Brumbaugh <benbrumbaugh@gmail.com>
Co-authored-by: Daisuke Suzuki <dai.suzuki.829@gmail.com>
Co-authored-by: xrbs00 <178640517+xrbs00@users.noreply.github.com>
2026-09-15 12:33:45 -05:00
brooklyn! 43bd5db429 fix(state): bound legacy transcript hydration memory
Co-authored-by: Benjamin Brumbaugh <benbrumbaugh@gmail.com>
2026-09-15 12:33:45 -05:00
teknium1 d84ece48b8 fix(mcp): Figma OAuth login completes despite the omitted iss parameter
Figma's authorization-server metadata advertises
authorization_response_iss_parameter_supported and its redirect omits iss,
so the mcp SDK's RFC 9207 check discarded every valid code and login never
finished. For that one issuer the provider fills a missing iss with the
discovered issuer and warns; a mismatching iss still fails and every other
server keeps the strict rule.

Fixes #111135
2026-09-15 09:29:00 -07:00
kshitijk4poor 5341f135a5 fix(update): the purge keeps hermes_constants; it is refreshed in place instead
hermes_constants owns the _HERMES_HOME_OVERRIDE ContextVar. Evicting it hands later
imports a fresh var: a reset token taken through the old module raises, and an
override set before the purge silently disappears. _reload_updated_runtime_modules
already re-executes it in place before every purge, so new symbols still arrive.
Caught by tests/hermes_cli/test_update_config_reload_tools_config.py on the stack.
2026-09-15 21:57:59 +05:30
kshitijk4poor d6ea7aa002 refactor(update): inline the tests exclusion; fix comments the wider purge made stale
`_STALE_PURGE_EXCLUDED_TOP_LEVEL` had one reader and a re-export nothing patched;
the config-check comment claimed root modules stay cached across the purge (no
longer true); the purge test module docstring still described package prefixes.
2026-09-15 21:57:59 +05:30
kshitijk4poor 9395f2f0d4 refactor(update): the purge scan has no fallback; trim to two invariants
The checkout root is where the update just pulled into, so "root unreadable" cannot
happen after a successful pull — drop the OSError fallback and the static five-name
tuple it fell back to (the tuple was the drift that caused the bug). Drop the phantom
`hermes_cli.hermes_logging` protection entry: no such module exists; the real
`hermes_logging` is root-level and now protected by name. Keep two tests: the stale
root `utils` scenario (red on base) and the hermes_logging protection.
2026-09-15 21:57:59 +05:30
Hubert Nimitanakit 29758a4eb2 fix(cli): purge every top-level checkout module, not just five packages
`_STALE_PURGE_PREFIXES` listed five package names, so every top-level module in
the checkout root survived the post-pull purge. `hermes update` then imported
new source against a cached pre-pull `utils`, `hermes_constants`, `plugins` or
`providers`.

Field failure (2026-09-12, macOS): updating 0.20.6 -> 0.21.2 crossed 3145986c20,
which added `base_url_origin` to `utils.py`. The restart phase's up-front
`from hermes_cli.gateway import ...` pulled `agent.auxiliary_client`, whose
`from utils import base_url_origin` hit the cached 0.20.6 `utils`:

    Update incomplete - gateway auto-restart failed: cannot import name
    'base_url_origin' from 'utils' (.../hermes-agent/utils.py)

The purge docstring already claims it evicts EVERY cached Hermes module; the
hardcoded tuple was the same "re-fixed per symptom" shape it replaced. Scan
`PROJECT_ROOT` instead: top-level `.py` files plus directories with an
`__init__.py`. 50 names here, and a newly added module can no longer drift out.

Two exclusions, both deliberate:

- `hermes_logging` joins `_STALE_PURGE_PROTECTED`. Its queue listener, handler
  list and `_logging_initialized` flag are module globals, so a fresh copy
  starts a second QueueListener over the same log files while the first runs.
- `tests` is never purged. pytest resolves fixtures through the identity of its
  already-imported test modules.

Falls back to the old tuple when the root is unreadable.
2026-09-15 21:57:59 +05:30
kshitijk4poor 4a2927d6d4 chore(contributors): map hcnimi for the #109314 salvage 2026-09-15 21:57:59 +05:30
teknium1 69fd61b0ef fix(desktop): anchor the first-paint backfill so long sessions stop lurching (#99920)
Switching to a long tool-heavy session still painted the tail, then jumped a
full viewport when the first backfill step committed, on current main.

Root cause: anchorBeforePrepend() skipped while the load was unsettled, but the
settle loop hands a bottom-pinned load back at the FIRST-PAINT height, before
the backfill transition commits. Every backfill step therefore ran unanchored;
the prepend grew scrollHeight by thousands of px with scrollTop untouched and
the view sat near the top until use-stick-to-bottom's ResizeObserver re-pinned
it frames later (measured CLS 0.6-1.0 per switch, layout-shift entries of 0.61).

Fix:
- anchor a bottom-pinned load even while unsettled (shouldAnchorBeforePrepend);
  an unsettled offset restore still never anchors, the settle loop owns it
- record the anchor's scrollHeight and consume it only on a taller tree, so a
  commit that does not grow content (transcript arriving under the first-paint
  budget, a store-window refresh) cannot spend the anchor as a no-op
- clear any anchor on a cold switch, not only on a key change: the stale rows
  shown under the new key are collapsed by the swap and the anchor with them
- do not arm the backfill while the transcript is empty or nothing is hidden

Live A/B (isolated Desktop, Xvfb, CDP layout-shift capture, 3 rounds x 2 long
sessions): main CLS 0.19-0.65 with one-frame 3-4k px scroll jumps every switch;
fixed CLS 0 on every warm switch and no scroll jump (cold first open of a
session keeps only the sub-0.1 first-paint shifts). Show-earlier and remembered
offset restore controls unchanged.
2026-09-15 08:51:14 -07:00
teknium1 043632e283 feat(desktop): hideable profile rail with a statusbar profile dropdown stand-in
For people who run profiles as bots, the colored profile strip at the sidebar
foot duplicates the sessions list (community request). Add a persisted
`hermes.desktop.profileRailVisible` preference (on by default) toggled from the
Sessions view menu ("Profile rail"), the shell right-click menu, ⌘K
("Toggle profile rail") and an unbound `view.toggleProfileRail` keybind.

While the rail is hidden the statusbar grows a `ProfileSwitcher` dropdown
beside the gateway switcher ("This device ⌄ · Profiles ⌄") offering the same
choices the rail does: this gateway's profiles, All profiles, every other
gateway's agents in fleet mode, New / Import / Manage. It also answers the
`profile.create` hotkey the rail used to own, so no door is lost.
2026-09-15 08:50:18 -07:00
teknium1 6bc3628ce8 fix(desktop): drag gateway/profile groups by their header, not only the hidden handle
The Gateway & profile sidebar sections carried dnd-kit listeners on the lead
glyph alone, and that glyph only reveals its grabber on hover — so a press on
the row's name or empty space did nothing and users fell back to the ⋯ menu's
Move up / Move down. Bind the sortable to the whole header (same shape as a
project row), with the handle and the ⋯/caret cluster keeping their own
gestures; a sub-threshold press on the label is still the fold click.

Live: Playwright pointer drag on the header label reorders sections (before:
unchanged, after: reordered); handle-drag still works. Test drives dnd-kit's
keyboard sensor on the label and is red on the old binding.
2026-09-15 08:50:18 -07:00
kshitijk4poor 78d338b9ee test(cli): drop the tautological mcp_servers assertion from the seed test
`_config_mcp_servers` is `self.config.get("mcp_servers") or {}`; with the
test's own `{"mcp_servers": {}}` it is `{}` whether or not the import fix
is present, so the assertion discriminated nothing. One invariant per fix.
2026-09-15 20:47:06 +05:30
mr-r0b0t 3d2842d84f fix(cli): import file_signature in TUI run-state init
#111408 widened the MCP config watcher seed from mtime to
utils.file_signature but omitted the import in cli_tui_mixin.
The NameError only fires when config.yaml exists, so isolated-home
tests short-circuited past it and every real CLI launch crashed.

Signed-off-by: mr-r0b0t <adam.manning@gmail.com>
2026-09-15 20:47:06 +05:30
kshitijk4poor 95dba8d9a5 fix(free-tier): classify_mint_exception keeps an uncoded AuthError's wait hint
The pure rewrite re-raises an uncoded `AuthError` as a server-error twin but
dropped its `retry_after`, so the mint memo and a sign-in `Failed` fell back
to the ladder / default wait instead of the wait the raiser named.
2026-09-15 20:44:42 +05:30
kshitijk4poor 658f319147 fix(free-tier): setup.ready carries the failure block flat, the shape setup.status already spreads
`SetupRecord.as_payload()` serialised the record verbatim, so the broadcast
nested `failure: {...}` while `setup.status` spread the same four keys flat.
A client keyed on `error_code` saw it on one surface and not the other. Flatten
it in `as_payload`, declare the three optional keys on `SetupReadyPayload`, and
regenerate the TS/OpenRPC contract.
2026-09-15 20:44:42 +05:30
kshitijk4poor 8f391f7d48 test(free-tier): reset the boot record and mint memo around each free_tier RPC test
`free_tier.provision` routes through the boot record when one exists, so a
has_identity record left by another file makes it skip the mint and the
lifecycle test fails whenever the two files share a process (CI's per-file
runner hid it). Reset both process memos before and after each test.
2026-09-15 20:44:42 +05:30
kshitijk4poor 683ca44046 fix(classifier): a welcome-host 403 naming the free tier itself is the tier refusing, not a billing wall
The "says something else" guard reuses the billing table, which carries the
Nous gateway's own free-tier phrases ("not available on the free tier",
"model_not_supported_on_free_tier"). On the welcome route those words mean
exactly "the tier refused"; routing them to billing prints a credits check to
an anonymous session that has none. Leave those two phrases on the
tier_disabled path.
2026-09-15 20:44:42 +05:30
Robin Fernandes 1034215ae8 docs(free-tier): drop the rehearsal page and its references
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 2a94ca80e7 fix(free-tier): review round 2 — route-gate the allowance verdict, keep policy/billing 403s, pool the provision RPC, guard the retry race
Should-fix
- _is_genuine_nous_rate_limit: the structured rate_limited verdict counts only
  on the welcome host; a paid-host 429 keeps main's exhausted-bucket rule.
- _nous_welcome_tier: the route-keyed dark-tier 403 applies only to a 403 that
  matches neither the content-policy nor the billing patterns, so a safety
  refusal or billing wall on the welcome host keeps its own recovery.
- free_tier.provision joins _LONG_HANDLERS (a forced mint + lock waits +
  re-inventory no longer block the RPC reader).
- retry_bootstrap_mint: under the lock, a build that found no identity never
  overwrites a record that has one (the loop racing the user's click).

Simplifications from the review
- _raise_for_anon_status is a (status, error) table; retryable derives from
  ANON_TERMINAL_CODES once (a bare 401 on sign-up now rides the ladder
  instead of dying for the process).
- classify_mint_exception is public and pure; the hand-built failure dict in
  free_tier.provision is gone (the memo is the one source).
- SetupRecord carries the memo payload as one `failure` dict instead of three
  unpacked fields.
- _welcome_surface_kind is a closed table with a "refused" default;
  _welcome_outage_copy excludes the classifier's `unknown` catch-all.
- FREE_TIER_RATE_LIMIT_CHAT is CARD + the sign-in tail, not a slice.
- Copy tests assert the contract (model named, tail present/absent) instead
  of freezing whole sentences.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes a241f42fc8 copy(free-tier): "it's free" without "keeps the free model" — signed-in free models are not the same model
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 2e3b8ecc82 fix(free-tier): the desktop card body leaves the "To sign in" tail off; the button is the door
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 16def7b8cc fix(free-tier): the desktop renders a free-tier refusal as its own card, not an OAuth re-login
A welcome-tier 403 classifies as auth_permanent, so the desktop's error
surface mapped it to "Your Nous Portal sign-in expired" with a Nous Portal
re-login button — the chat sentence never reached the user. Terminal results
on the free route now carry a structured free_tier block (kind + the chat
sentence); agent/error_surface.py turns it into a free_tier_<kind> code on
the provider layer with the sentence as `message`. The desktop gives those
codes their own titles, shows the backend sentence as the body, and offers
"Sign in with a Nous account" (the free-tier dialog) instead of the OAuth
re-login, with Retry only where a later send can succeed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes d89cacc25f chore(free-tier): keep the rehearsal server out of the repo; the docs page explains the stand-in instead
The fault-injecting server served one-off manual rehearsal only and would
drift silently from the real services; the doc now says how to point the
desktop at any local stand-in (the three env overrides) and what such a
stand-in has to speak. The dev-only HERMES_EXTRA_WELCOME_HOSTS override stays,
pinned by a test in test_anon_failure_modes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 59fad62a40 fix(free-tier): review follow-ups — read the classifier's context, never replace a locked identity, re-inventory on retry
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
  the long-wait rate-limit check read the turn's extract_api_error_context()
  dict, which never carries welcome_refusal / welcome_route. They now read
  classified.error_context, where _nous_welcome_tier parks them; the guard
  records the classifier's reset_at. Tests drive the real classifier and the
  real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
  account (anon_account_locked) is now retired without replacement, matching
  the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
  re-inventories, so a provider connected during the cooldown keeps
  inference.
- The desktop's setup.ready listener only refreshes an untouched picker
  (oauth mode, no local endpoint, idle flow) and re-checks after the
  readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
  lock that _send re-acquires; the log is copied out first.

Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
  behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
  scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
  extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
  loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 51e39af967 feat(free-tier): ruled behaviour for every welcome-api failure, with friendly copy and a fault-injecting rehearsal server
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.

Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
  temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
  into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
  honours the server's wait, climbs a short ladder when the service is
  unreachable, never retries terminal codes, and yields to the user's own
  retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
  background loop retries transient failures and re-announces setup.ready.
  setup.status and free_tier.status expose the block; free_tier.provision is
  the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
  the route); model_not_free moves onto the gateway's alternate once;
  anon_on_paid_host re-reads the route once; a long rate_limited refusal
  trips the cross-session guard; a locked account is retired but never
  replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
  retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
  service is off" (what is unavailable is using Hermes without signing in,
  and signing in is free), no jargon, spoken waits.

Desktop
- A setup-failure notice above the provider picker: one sentence per code,
  a retry when the backend says one can work, the sign-in pointer only when
  the account service answered at all. The overlay re-checks readiness on
  setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.

Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
  real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
  (dev-only, env-only) lets the route rules treat it as the welcome host.
  Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
teknium1 24fd22b94d test(cli): trim #110737 re-queue coverage to two invariants on the mixin
Drive CLIChatTurnMixin._chat_render_turn through a minimal stub instead of a full
HermesCLI (prompt_toolkit stubs, cli reload): the invariants are that a (text, images)
payload reaches _pending_input intact and that several queued parts join their text
and merge their images. Both are red on origin/main.
2026-09-15 06:36:55 -07:00
Kevin Rajan 40b22806f8 fix(cli): preserve image attachments when re-queueing interrupt messages
_clui mixin bundles Enter submissions carrying images as (text, images)
tuples, and _tui_enter_while_busy puts those tuples on _interrupt_queue in
interrupt mode. _chat_render_turn then did "\n".join(all_parts) over the raw
payloads, raising TypeError on the tuple — swallowed by the outer chat()
handler, so the interrupt message was silently lost and no next turn started.

Unpack tuple payloads in the re-queue block: join the text parts and re-attach
all images as a single (text, images) tuple, matching the shape
_tui_process_one_input already accepts. Plain-string payloads are unchanged.

authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction; all changes reviewed and approved by the contributor
2026-09-15 06:36:55 -07:00
teknium1 64aeb5dd1a fix: gate the stalled-redraw close on socket identity, not session.attached
When viewer A's replay fails because viewer B superseded it mid-send, attach(A)
returns False while session.attached is True for B. The handler then logged a
misleading "pty input stalled … recycling terminal session" warning and issued
a second close on A. Checking `session._ws is ws` limits the close to the one
case it exists for: this socket is still attached and its redraw write stalled.

Review finding: superseded socket's failed replay logged a stall and re-closed itself because the gate tested session-wide state.
2026-09-15 06:36:19 -07:00
teknium1 7a1d278528 fix(dashboard): skip the stalled-input close when attach detached itself
attach() now returns False both when the client dropped mid-replay (session
already detached, socket dead) and when the force-redraw write stalled
(socket still attached, worth closing with 1013). Only close the socket in
the second case; the registry detach stays as the guard for both.

Also trims the salvaged tests to two invariants, both red on origin/main:
a mid-replay drop leaves the session reap_idle()-reclaimable, and a drain
send failure detaches the current socket without touching a replacement
that attached during the send (#110849).
2026-09-15 06:36:19 -07:00
MKanso 941d8b59a7 fix(pty): handle replay send failures cleanly 2026-09-15 06:36:19 -07:00
MKanso 6d622d8eba fix(pty): recover sessions after socket send failures 2026-09-15 06:36:19 -07:00
teknium1 9c311038fa fix(discord): slash registration and /skill refresh scan the catalog off the event loop
connect() called _register_slash_commands inline on every (re)connect, and
/reload-skills called refresh_skill_group inline; both run
discord_skill_commands_by_category, the same per-skill path-resolution walk
the Telegram menu paid in #110707. On a 1.5k-skill install that holds the
loop past the liveness watchdog. Registration now hops through
asyncio.to_thread from connect(); refresh_skill_group is a coroutine that
hops the rescan the same way (the reload handler already awaits an awaitable
result). Contextvar-scoped profile overrides propagate through to_thread.

One invariant test: the loop keeps ticking while the scan blocks, from both
sites. Red on the PR head, green here.

Review finding: Discord _register_slash_commands/refresh_skill_group ran the skill catalog disk scan synchronously on the event loop.
2026-09-15 06:35:04 -07:00
teknium1 ab782a7685 fix(gateway): skill-slash fallthrough and the Telegram inline picker run off the event loop
The gateway's idle-command path resolved skill slash commands inline on the
loop: a cold skill scan, skill file loads and the unavailable-skill rglob over
every skills dir. On a 1.5k-skill install that held the loop ~2 minutes, the
loop-liveness watchdog fired and the gateway exited mid-session (#111091).
_hm_skill_slash_rewrite now runs through _run_in_executor_with_context so the
profile contextvars the scan is scoped to survive the hop. Known commands still
short-circuit before any I/O (previous commit).

Same class in the Telegram inline picker: build_inline_results rebuilds the
command/skill catalog per keystroke via _collect_gateway_skill_entries, the
same path-resolution pass #110707 traced in the command menu.

One invariant test: a known command triggers no scan; an unknown command's scan
runs while the loop keeps ticking. Red on origin/main and on the reorder-only
tree, green here.
2026-09-15 06:35:04 -07:00
KoNit-K 88c04e655f fix(gateway): known slash commands never pay the unavailable-skill scan
Reorder _hm_unknown_slash_reply so the I/O-free GATEWAY_KNOWN_COMMANDS
membership test short-circuits before _check_unavailable_skill walks the
skills tree (#111091). Taken from #111099 by @KoNit-K; the index cache half of
that PR is superseded by running the scan off the loop.
2026-09-15 06:35:04 -07:00
teknium1 660d4b8d87 fix(telegram): hop only the menu build off the loop; one invariant test covers both menu sites
Follow-up to the salvaged #110716 commit. Drops the module-level
_build_telegram_command_menu wrapper (telegram_menu_max_commands is a config
read and stays on the loop; asyncio.to_thread takes kwargs directly), and
applies the same hop to _ensure_forum_commands, which rebuilds the menu on the
inbound-message path for every forum chat until registration succeeds — the
second live fire site named in #110707.

Replaces the contributor's test with one that proves the invariant for both
sites: the loop keeps ticking while telegram_menu_commands blocks. Red on
origin/main, green here.
2026-09-15 06:35:04 -07:00
KoNit-K 0a5846d37f fix(telegram): build command menu off event loop 2026-09-15 06:35:04 -07:00
teknium1 14bd358f5e fix: unlink the dm plaintext file too once a live delivery settled
The live branch of _run_delivery returns from _wait_live_dm before the
try/finally that removes <dm>.txt, so a settled delivery dropped its
.live.json intent but left the sibling .txt — the same plaintext — on disk.
_wait_live_dm now takes the dm file and removes both on settled; pending
and failed outcomes still keep both for the retry.

Review finding: settled live path unlinked <dm>.live.json but the sibling <dm>.txt survived.
2026-09-15 06:34:25 -07:00
John Paul Soliva 1238dbc09d fix(bot-mode): live-delivery intents are swept, and dropped once the owner settled
A live DM's intent file (<dm file>.live.json — owner, delivery id and the
message plaintext) has to outlive its runner so a retry replays the same
delivery id instead of minting a second one. Nothing ever removed it: the
runner unlinks only the dm file, and cleanup_bot_dm_cache sweeps only *.txt,
so every live DM left its plaintext in the DM cache directory indefinitely.
Now the sweep reaps intents past the stale cutoff, and a delivery the owner
settled drops its intent immediately — nothing retries a settled delivery.
2026-09-15 06:34:25 -07:00
teknium1 15ebf7ec90 test(desktop): drop the never-assigned descriptor-profile knob from the connections mock
The salvaged switch-back regression only ever exercised the profile-less
primary shape; the mock's string branch had no writer. Publish the
profile-less local descriptor unconditionally and say why.
2026-09-15 06:32:53 -07:00
Konstantin Khlopkov 825324fdc9 fix(desktop): restore the last-used profile when the primary local descriptor is profile-less
The local primary backend (startHermes) publishes its descriptor without a
profile field, unlike the pooled/forced-local child path. Downstream, a
profile-less descriptor normalizes to 'default', so on a switch back to the
local source the store either refused to remember the source/profile pair or
killed the commit in targetIsActive() — landing the window on 'default'
every time (#110819).

- The primary local descriptor now carries the profile it booted with, and
  the unshared primary branch of ensureBackend() keeps the requested key on
  the descriptor instead of dropping it.
- selectConnection() tolerates a profile-less descriptor on the source it is
  landing: the activation already published the route that was asked for, so
  a missing profile no longer strands the switch.
2026-09-15 06:32:53 -07:00
teknium1 c30b3c1a2c refactor(web): drop the dead shouldPinScroll helper and trim the reveal tests to two invariants
ChatPage no longer pins the page to (0, 0); keyboardRevealScrollDelta is the only caller-facing
helper left in keyboard-inset.ts. Update the isActive-effect comment that still described the pin.
2026-09-15 06:32:12 -07:00
edrethardo e72813393c fix(web): jump the iOS viewport to the chat composer when the keyboard opens
Pinning the dashboard to (0, 0) on keyboard inset fights Safari's visual
viewport and leaves the Ink input line off-screen. Scroll by the delta
between the xterm host bottom and visualViewport bottom instead, and
re-sync after focus while the keyboard animates.

Fixes #110414
2026-09-15 06:32:12 -07:00
teknium1 b3feb88a95 fix(tui): skip the session-store title read when the Bot Chat hint is set
_is_bot_mode_session runs on every completion of every session; a non-empty
title hint already decides the answer, so only an empty hint should fall
through to the live-title lookup.

Review finding: plain sessions paid a session-store read per turn in _is_bot_mode_session.
2026-09-15 06:31:31 -07:00