Commit Graph

28271 Commits

Author SHA1 Message Date
Kshitij Kapoor 72ae855da5 refactor: consolidate _db_persisted literal + marker-insensitive no-op check (#92231 follow-up)
Review-pass follow-up on the load-time durability stamp:

- hermes_state.py: import the marker from agent.context_compressor instead
  of a third synced literal (hermes_state already imports agent.* at module
  level; only run_agent is circular). Old comment claimed otherwise.
- agent/turn_finalizer.py: replace the raw "_db_persisted" string at the
  fill-empty-tail pop site with the shared constant (was outside the drift
  guard).
- agent/conversation_compression.py: the no-op progress check now falls back
  to a marker-insensitive comparison (_strip_marker_for_comparison). Loaded
  rows are stamped at materialization time while compress() output is
  marker-swept, so a semantically-identical no-op copy on a cold-resumed
  session would previously compare unequal and take the progress branch.
  Raw == still runs first so engine-returned list subclasses keep their
  __eq__ semantics.
- test_marker_constant_in_sync extended to turn_finalizer + identity
  assertions; new test_noop_progress_check_is_marker_insensitive
  (mutation-checked: fails when the helper is neutered).
2026-08-23 13:34:32 +05:30
kshitijk4poor 016ba66176 fix: stamp _db_persisted at row load time so resumed transcripts never re-append (#92231)
Resumed sessions loaded message dicts from state.db WITHOUT the
_DB_PERSISTED_MARKER, so any flush that lost the identity boundary
(compression durable-snapshot adoption, incremental tool-call persists,
rotation preflight on cold resume) re-appended the ENTIRE loaded
transcript as new rows. Compression cycles then doubled the copies:
the incident session grew 998 -> 1995 -> 3990 -> 7981 rows across
three aborted rotations (15,962 active rows, only 472 distinct).

Fix at the architectural chokepoint: SessionDB._rows_to_conversation
(shared by get_messages_as_conversation and get_resume_conversations)
now stamps the marker at row materialization time - a dict built FROM
a durable row is persisted by construction, regardless of which caller
loads it or how the list is later handed to a flush.

Safety:
- Wire-safe: every transport strips underscore-prefixed keys before
  the API request (chat_completion_helpers, anthropic_adapter), same
  contract as the existing _row_id stamp in the same function.
- Rotation handoffs still write: compression's assembly copies strip
  the marker (_fresh_compaction_message_copy + the terminal
  _strip_persistence_markers sweep), so compacted transcripts still
  flush to the child session (#57491 invariant preserved).
- Branch/seed copies unaffected: /branch and _persist_branch_seed
  build fresh field-projected dicts and write via append_messages_batch
  directly, not through the marker-gated flush.

Tests: new regression suite (marker sync, load stamping, 3-cycle
amplification repro, new-tail write guard, compaction-copy handoff);
updated the #68454 control test that asserted the old double-write
behavior and the ACP restore shape test.
2026-08-23 13:34:32 +05:30
kshitijk4poor dff84f1890 fix(browser): cap browser_vision native embeds for history reuse
browser_vision's native fast path base64-encoded screenshots at full
resolution and baked them into the tool result uncapped — the exact
sibling of the vision_analyze path #92699 fixed. Apply the same
proactive 256KB/1568px resize before the embed enters reusable history.

Fail-open by design: without Pillow the resize helper falls back to raw
bytes and the compressor's keep-newest pass still retires stale embeds.

Sibling-gap follow-up for the #92725 salvage; the shared-cap approach
mirrors the policy-owner idea from #92748.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-08-23 13:27:04 +05:30
HexLab98 7ff2fe8bc9 fix(compression): retire stale vision tool images in the protected tail
Images locked in protect_last_n never shrank, so compression savings
stayed under 10% and anti-thrash disabled further compaction. Keep the
newest three tool-result screenshots live for follow-up QA and replace
older native embeds with placeholders.
2026-08-23 13:27:04 +05:30
HexLab98 21a93f0a67 fix(vision): size native embeds for history reuse
vision_analyze baked up to 4 MB / 7900px screenshots into immutable
history, so every later turn re-sent ~400K chars. Cap embeds at 256 KB
and 1568px (the long edge models actually read) so screenshot QA no
longer blows the context.
2026-08-23 13:27:04 +05:30
Teknium 3c44cd0c67 docs: record why the reap grace window exists in the reaper docstring
Follow-up for salvaged PR #91994 -- without this note a future cleanup
pass could read the age gate as dead weight.
2026-08-23 00:52:50 -07:00
Ishman82 b44c2bdab2 fix(desktop): spare concurrently starting backends
Desktop writes backend.lock.json only after a remote profile backend announces readiness. Concurrent Bot Mode profile starts could therefore see their young siblings as unowned PPID-1 processes and mutually reap them, causing SSH reconnect storms and stale turn leases.\n\nProtect unregistered backends for a bounded startup grace period, fail closed when age cannot be read, and cover young, unknown-age, boundary, lock-owned, and old-orphan behavior.
2026-08-23 00:52:50 -07:00
Teknium 8804e78354 feat(dashboard): Desktop and dashboard read the update receipt instead of inferring success (#91277 Phase-1 bullet 3)
Builds on @mrsucesso's durable-marker recovery (previous commit):

- GET /api/hermes/update/receipt — the full durable receipt (steps,
  skips, gateway restart outcome, fleet matrix) + compact summary; the
  authoritative update-outcome record (written by every run since
  #91283, including refused/failed).
- /api/actions/hermes-update/status now attaches the receipt summary,
  and when BOTH the in-memory registries and the update.log marker are
  gone (dashboard restarted + log rotated — the #81193 state), a
  finished receipt reports the outcome: success→0, partial→1. A
  still-running receipt proves nothing (clients keep polling).
- Desktop (updates.ts): the apply poll reads the attached receipt — a
  finished receipt whose run started at/after this apply is
  authoritative, replacing timeout-based failure inference across the
  update's restart gap ('Backend update failed' on successful updates,
  #81193; 'boot failed' during update restarts, #87359).

Live-verified: real uvicorn server + real UpdateReceipt writer (the
exact code hermes update runs) over real HTTP — receipt endpoint 200
with summary; #81193 state (no registries, no marker) reports success
from the receipt alone; partial receipt with a DOWN fleet row maps to
exit 1 (no false success).
2026-08-23 00:50:41 -07:00
Mauricio Ruiz 091106092b fix(dashboard): recover update success after restart
Persisted update completion markers survive the dashboard restart that clears in-memory action state. Recover the latest safe marker from update.log so remote Desktop clients do not report a successful backend update as failed.
2026-08-23 00:50:41 -07:00
Teknium 4553e71993 fix(test): isolate the fork-sync test from the host machine; abort at the reload proof point
The salvaged test drove the FULL post-update pipeline against the real
dev box: real fleet probes read live gateways as STALE (exit 1) and the
restart phase tripped the live-system guard on a real gateway PID. Pin
empty fleet/gateway discovery and make _reload_updated_runtime_modules
(the proof the post-update path ran — the bug returned before it) abort
the pipeline. Sabotage-verified: disabling the hoisted sync fails the
test.
2026-08-23 00:19:46 -07:00
Franci Penov e366df6889 fix(cli): treat a fork's upstream sync as an update
On a fork, `hermes update` compares HEAD against origin/main, and only then
syncs the fork from upstream — inside the `commit_count == 0` branch, which
returns immediately afterwards. So an update that pulls hundreds of commits
from upstream prints "Already up to date!" and skips everything the
post-update path does, including the dependency sync and the gateway restart.

Observed on a fork-based deployment: 1654 commits pulled, "Already up to
date!", and the launchd gateway left running. It then held pre-update modules
in memory while lazily importing post-update ones, and failed later with an
AttributeError for a method that plainly exists on disk — a mixed runtime that
looks nothing like an update problem. Correlating every run in update.log, a
restart happened on exactly the runs that pulled upstream *without* also
claiming to be up to date, and never once they started co-occurring.

Decide before the branch: capture HEAD, sync, and if HEAD moved, set
commit_count from the range so the normal post-update path runs. The pull that
follows is a no-op (the sync updates origin too); reaching the restart is the
point. commit_count is floored at 1 — HEAD moving *is* the update, so a failed
or zero count query must not send us back down the early return.

steps still being skipped afterwards.

Refs #73108
2026-08-23 00:19:46 -07:00
Teknium 1684877868 fix(update): a gateway killed by the restart phase and never replaced now fails the fleet check (DOWN row)
Phase-1 verification gap (#91277, found auditing our own landed matrix
against the mapped issues): collect_fleet_versions only listed gateways
with a LIVE pid, so 'restart stopped it and nothing came back' produced
NO row at all — the exact silent-failure shape the matrix exists to
catch (#88848/#74973 class) passed with exit 0.

- collect_fleet_versions(pre_restart_pids=...): a dead pid becomes a
  'down' row only when it was alive at update start AND its runtime
  status still claims a running state. Rollout-safe: no snapshot (old
  callers), clean stops, startup failures, and stale records from
  long-dead gateways keep the historical no-row behavior.
- print_fleet_version_matrix escalates on down rows like stale ones
  (exit 1) with the per-profile restart remediation.
- cmd_update passes its existing pre-restart PID snapshot.

Sabotage-verified (reverting the membership check fails the new test);
live-verified with a real spawned-then-killed process producing the
DOWN row and matrix escalation.
2026-08-22 23:46:06 -07:00
Teknium 32a8a7031e chore: map contributor email for salvage attribution 2026-08-22 23:25:43 -07:00
Josh Holt 5f0a8f8739 fix(picker): harden keyless provider gate logging and credentials path validation 2026-08-22 23:25:43 -07:00
Josh Holt b9f17ba3f1 fix(picker): scope Vertex explicit-config to Hermes signals, not ambient ADC
Address review feedback: the gate reused has_vertex_credentials(), which also
returns True for an ambient GOOGLE_APPLICATION_CREDENTIALS path. That var is
commonly set globally for unrelated GCP work, so a user who never configured
Hermes for Vertex would see it in the explicit-only picker and could spend
against those credentials — weakening the explicit-configuration guarantee the
gate documents (mirrors the existing _IMPLICIT_ENV_VARS carve-out).

Add has_explicit_vertex_config() in agent/vertex_adapter.py that checks only
Hermes-scoped signals — VERTEX_PROJECT_ID / vertex.project_id (project
override) or a resolvable VERTEX_CREDENTIALS_PATH — and NOT
GOOGLE_APPLICATION_CREDENTIALS. Route the auth gate through it.

Adds a regression test asserting an ambient GOOGLE_APPLICATION_CREDENTIALS
path alone does not mark Vertex explicit, and updates the existing test to
drive the real config signal instead of mocking has_vertex_credentials().
2026-08-22 23:25:43 -07:00
Ajad van Wyk 3503c06d80 fix: surface Bedrock in explicit-only model pickers when AWS env credentials are set
is_provider_explicitly_configured() only checked provider env vars for
auth_type="api_key" providers. Bedrock is registered with
auth_type="aws_sdk" and an empty api_key_env_vars tuple, so a user who
sets AWS_BEARER_TOKEN_BEDROCK (or an AWS_ACCESS_KEY_ID +
AWS_SECRET_ACCESS_KEY pair) in .env was never counted as having
explicitly configured the provider.

Symptom: the desktop model picker (and any consumer of
build_models_payload(explicit_only=True)) silently hid the Bedrock row
even though list_authenticated_providers had discovered credentials and
built a full model list for it. Reproduced on main:

    build_models_payload(ctx, explicit_only=False)
      -> ['moa', 'nous', 'bedrock', ...]        # row exists, 132 models
    build_models_payload(ctx, explicit_only=True)
      -> ['nous', ...]                          # bedrock filtered out

Fix: aws_sdk-type providers now count as explicitly configured when
Bedrock-relevant env credentials are present. Deliberately env-var-only:
ambient sources (AWS_PROFILE / SSO, EC2 IMDS, container credentials)
still do NOT auto-surface, consistent with the gate's purpose (#56974)
and with a lone AWS_ACCESS_KEY_ID (no secret) not counting.

Tests: six behavior-contract cases in test_auth_provider_gate.py
covering bearer token, key pair, lone key id, ambient AWS_PROFILE,
no-credential baseline, and non-leakage into other providers.
2026-08-22 23:25:43 -07:00
Kevin Yin 9ce46a0968 fix(models): bind Anthropic pool key to its endpoint 2026-08-22 23:25:29 -07:00
Kevin Yin 8ede2e1472 fix(models): discover Anthropic pool API keys 2026-08-22 23:25:29 -07:00
Teknium e132e11ea7 Merge pull request #92731 from NousResearch/salv/90006-remote-bots
feat(desktop): remote bots open their own Bot Chat without re-homing Desktop (salvage #90006)
2026-08-22 23:25:19 -07:00
Teknium e6962b818c style: eslint --fix across src/ + electron/ (CI lints the full tree); drop unused destructure 2026-08-22 22:35:55 -07:00
Teknium 894d02191d Merge current main 2026-08-22 22:26:14 -07:00
Teknium 278dac43b1 chore: map contributor email for attribution gate 2026-08-22 22:26:08 -07:00
Teknium 4e4ee7ee75 style: eslint --fix on merged TS surfaces 2026-08-22 22:25:59 -07:00
Teknium 0404020f7b Merge PR #90006: connection-bound Bot Mode actions, reconciled with name-identity + fail-closed canonical resolution
Salvage of saralilyb's remote-bot routing work onto current main:
- kept: immutable (connectionId, profile) owner capture, requestForBot
  routing, backendTargetProfile aliasing, group session owners,
  connection-qualified deletion, focused-owner atoms, remote roster
  merge, Electron profile-delete routing, sdk/store/transcript changes
- reconciled: canonical Bot Chat resolution stays NAME-identity (the
  'Bot Chat' registry row) and FAIL-CLOSED on lookup errors — now
  consulted on the bot's own source via the captured owner route, so
  remote bots get the same no-fork guarantees
- dropped: pointer-pin plumbing (preferredSessionIds, saveBotMeta chat
  writes, pin verification) — superseded by name-identity on main;
  renderer-side remote DM delivery (deliverRemoteRosterMentions /
  pollRemoteDmReply / ensureRemoteCanonicalChat) — superseded by the
  message_agent tool architecture (#91802/#91915: middleware identifies,
  never delivers); pointer-era test files deleted on main
- openStoredBotChat/createCanonicalChat: remote opens keep Desktop's
  chrome home (keepAllProfilesScope: true on routed opens); local bots
  keep the measured workspace re-home
- prepareBotSource: capability gate only — routed RPCs never require
  activation authority
2026-08-22 22:24:54 -07:00
hermes-seaeye[bot] 8b09a9df84 fmt(js): npm run fix on merge (#92718)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-23 05:11:16 +00:00
hermes-seaeye[bot] d49d495c2b fmt(js): npm run fix on merge (#92714)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-23 05:06:17 +00:00
Andrex Ibiza, MBA 38ce2d7553 fix(desktop): enforce exact route identity authority
Make explicit registry qualification authoritative: only a current exact ID is accepted, while blank, malformed, unknown, or retired claims fail closed without endpoint inference. Restrict genuinely unqualified legacy descriptors to the shared full-envelope URL/Cloud/SSH matcher, reject zero or multiple matches, normalize SSH host/user identity, and prove remote-primary restoration keeps the exact (connectionId, profile) tuple.

Closes #90048.

Prior work by @teknium1 in #89719 and #88922, @andrexibiza in #90913, and @AndreasG78 in https://github.com/NousResearch/hermes-agent/issues/90048#issuecomment-5375227679 shaped this implementation. @saralilyb's #90006 remains downstream consumer context; the production stopgap is credited but excluded because registry primary does not prove route ownership.
2026-08-22 22:01:39 -07:00
Teknium 4f0e466e5d fix: accept pool-only Anthropic OAuth entries in the desktop picker filter
The salvaged carve-out covered the PKCE file and Claude Code credentials
but missed the canonical wired-token location — auth.json
credential_pool.anthropic oauth entries. Discovery accepts those via
pool.has_credentials(), so the filter must too. Read-only dict access;
api_key pool entries intentionally stay excluded (E2E case 4).
2026-08-22 22:00:19 -07:00
Teknium aa24d87801 chore: map contributor email for salvage attribution 2026-08-22 22:00:19 -07:00
Viktor 4ec57d56a9 fix(inventory): keep Anthropic OAuth logins visible in desktop pickers
The desktop explicit-only picker filter drops any provider row where
is_provider_explicitly_configured() is False. That gate only recognizes
active_provider, model.provider, and API-key env vars — so a user who
authenticated Anthropic via OAuth (Hermes device flow or Claude Code
~/.claude/.credentials.json) had the Anthropic row silently hidden from
the desktop model picker even though list_authenticated_providers()
accepted those same credentials when building it (model_switch.py has an
equivalent special-case in its discovery path).

Unlike ambient CLI tokens (gh -> copilot), an OAuth access token only
exists after an interactive login, so its presence is deliberate user
configuration. Add a narrow carve-out for the anthropic slug that keeps
the row when either OAuth reader finds an accessToken; the strict gate
still governs every other provider, and is_provider_explicitly_configured
itself stays untouched so PR #4210's consent guard for auxiliary tasks
keeps its current behavior.
2026-08-22 22:00:19 -07:00
Teknium 2a379a1675 chore: map salvage contributors (klaus765, ctaylor86) 2026-08-22 21:55:39 -07:00
Teknium c942cd9ea1 fix(desktop): settings scope requests can never target primary by accident
The settings 'Applies to' store uses null = 'follow the active profile',
but the API helpers (profileScoped/capabilityScoped) use null = 'target
the primary/default backend'. Every page that passed the raw override to
a request silently read/wrote the primary profile whenever no override
was set — writes landed on the right profile via other paths while reads
repainted primary's values, so profile model changes appeared to revert
(#90549 class).

Close the class at the seam instead of per call site:
- store/settings-scope: new $settingsRequestProfile computed — the
  request-shaped scope (string | undefined, never null). Documented as
  THE value to hand to API helpers.
- config-settings, keys-settings, messaging: consume the request-shaped
  computed; ModelSettings/MemoryConnect/ProviderConfigPanel/
  useEnvCredentials props narrowed to string | undefined so a
  primary-targeting null can no longer be plumbed through.
- keys-settings site was a live third instance: getEnvVars(null) read
  primary's env store on non-default profiles.

Regression tests: store computed shape, ModelSettings unscoped+scoped
reads, KeysSettings unscoped fetch (all fail against the old behavior;
sabotage-verified).
2026-08-22 21:55:39 -07:00
Carl Taylor abd7f75b8d fix(desktop): keep Messaging on active profile 2026-08-22 21:55:39 -07:00
Klaus Suppan 680b11503c fix(desktop): model settings follow active profile instead of primary
ModelSettings passed scopeProfile (null when following the active
profile) directly to the Hermes API helpers. The helpers interpret
null as "target the primary/default backend", not as "follow the
active profile". This meant that when a non-default profile was
active (e.g. local with LM Studio), the model picker showed the
primary/default profile's providers instead — LM Studio was
invisible.

Fix: convert null → undefined before calling the API helpers, so
profileScoped() falls back to the app-wide active profile.

Added regression test verifying the helpers receive undefined (not
null) when no scope override is set.
2026-08-22 21:55:39 -07:00
Teknium 8b86097a62 fix(gateway): honor env-configured local backends and persist Tool Gateway declines
Follow-ups on top of the salvaged #92665 commits, closing the two gaps
called out in #92647:

- SEARXNG_URL / CAMOFOX_URL now count as direct web/browser configuration
  in _get_gateway_direct_credentials(), so env-only keyless local setups
  are offered unchecked instead of pre-checked.
- Submitting the checklist with unconfigured tools left unchecked records
  them in tool_gateway_declined_tools (new known root key) and stops
  pre-checking them on subsequent Nous model swaps; opting in later clears
  the decline. Cancel (Ctrl-C/ESC) records nothing.
- has_direct labels mention SearXNG/Camofox when that's what's detected.
- Docs: tool-gateway.md gains an enablement-checklist section.
2026-08-22 21:36:55 -07:00
chelsealong f93593d165 test(gateway): cover browser-use as an explicit non-nous selection
Addresses maintainer follow-up on PR #92665: confirm _selected_provider
correctly excludes an explicit BYOK browser.cloud_provider: browser-use
selection from get_gateway_eligible_tools' unconfigured/has_direct
buckets, same protection already covered for web.backend: searxng.
2026-08-22 21:36:55 -07:00
chelsealong ce9ddd35e7 fix(gateway): fix 3-tuple/4-tuple arity crash in get_gateway_eligible_tools
The early "fail closed" returns (account fetch error, not entitled,
non-nous provider) still returned 3-element tuples after the function's
happy path and every caller moved to a 4-tuple, crashing
prompt_enable_tool_gateway with ValueError for any logged-in Nous
account that isn't paid or pool-entitled — the common case, hit
unconditionally from `hermes model`.
2026-08-22 21:36:55 -07:00
chelsealong 4edb24276f fix(tools): don't pre-check keyless local backends in Tool Gateway checklist
get_gateway_eligible_tools() classified a tool as "unconfigured" whenever
it found no direct API-key credential, ignoring an explicit non-nous
selection already stored in config (e.g. web.backend: searxng, a keyless
self-hosted backend). Every unconfigured tool is pre-checked in the
`hermes model` Tool Gateway checklist, so a single Enter during Nous model
setup silently rewrote web.backend (and similarly browser.cloud_provider,
tts/stt/image_gen.provider) to "nous" for users who had deliberately
configured a keyless local backend.

get_gateway_eligible_tools() now also resolves each tool's stored
selection via the existing _selected_provider() helper and routes an
explicit non-nous selection into a new explicit_configured bucket instead
of unconfigured. prompt_enable_tool_gateway() never offers those tools,
so they can no longer be pre-checked or accidentally overwritten.

Fixes #92647
2026-08-22 21:36:55 -07:00
Teknium f4067774aa fix(update): token-based control-plane classifier + live E2E for the Desktop-lifecycle cold-start skip (#76129 salvage follow-up)
On top of @686f6c61's premise-corrected #76745:

- _looks_like_desktop_control_plane now uses the parser-derived
  _hermes_holder_subcommand instead of substring matching — the
  #90778/#91869 class ('-m dashboard chat' and 'kanban --preserve-cache'
  argv no longer read as control planes). Regression test added,
  sabotage-verified (reverting to substrings fails it).
- Live E2E (this host, real processes + real spawn ledger): live
  supervised serve owns lifecycle; killed spawner (orphan) does not;
  dead serve entry excluded; empty ledger does not.
- Live Windows E2E for the wine2e lane: real self-registered ledger
  entry suppresses the actual cold-start plan; dead serve restores it;
  holder-scan fallback rung proves the token classifier live.

Co-authored-by: 686f6c61 <github@00b.tech>
2026-08-22 21:36:44 -07:00
686f6c61 4ccc4b6931 fix(update): skip Windows gateway cold-start when Desktop owns lifecycle
Vestigial autostart is not proof the user wants a standalone gateway
run. When Desktop currently supervises this install's control plane,
the updater must not spawn a competing messaging daemon. Serve is not
treated as gateway-equivalent.
2026-08-22 21:36:44 -07:00
Teknium 87b645f52c fix(desktop): a failed Bot Chat registry lookup no longer forks the bot's forever chat
findExistingCanonicalChat() swallowed every lookup error and returned
null — indistinguishable from 'this bot has no Bot Chat yet' — so a
transient RPC failure against a just-restarted backend (the exact
post-desktop-update window) sent createCanonicalChat() straight to
session.create, minting a fresh 'Bot Chat' while the real one (data
intact, hidden) still held the canonical title. Users experienced this
as bots losing all context after every desktop update.

The lookup now fails CLOSED: a failed registry consultation throws,
both open paths surface their existing 'try again' toast, and
session.create can never fire off an unknown ownership state.

Tests: two new VM-executed regression tests (sabotage-verified — both
fail with the old fail-open catch); hide-bot-chats source-shape regex
updated for the new layout. 364/364 plugin tests green.
2026-08-22 21:24:30 -07:00
Teknium c9c44d0df9 Merge pull request #92636 from NousResearch/fix/windows-launcher-managed-bin
fix(windows): stage hermes launchers in the managed binary dir, not the git checkout
2026-08-22 20:58:00 -07:00
emozilla 0b01599eef fix(tests): repair the two Linux-lane CI failures on this branch
Both failed only in the full Linux suite, which the targeted local
battery never ran:

- test_update_zip_two_phase.py's AST guard (#76105) flags any code
  literal "Scripts" in hermes_cli as an open-coded venv layout.
  migrate_windows_bin_path's legacy PATH key now derives it via
  venv_bin_dir(root / "venv", windows=True) — same value, canonical
  helper. The literal `venv` component stays: the key must match what
  the pre-#83797 installer wrote to the registry, not where the venv
  lives now.
- The managed-bin marker tests built expected PATH entries from
  tmp_path, so on a POSIX host they compared forward-slash strings
  against the backslash markers and could never match. Markers match
  Windows registry PATH entries, so the tests now feed Windows-shaped
  literals — host-independent, same contract.
2026-08-22 23:10:32 -04:00
Teknium fd760435c6 test(bedrock): make the botocore stub windows airtight — kills the vendored-import flake
CI flake mechanism (PR #92617 red, reproduced standalone): tests plant
fake botocore modules via patch.dict; when the REAL botocore.exceptions
is first imported in an interpreter state where a fake parent is (or
was) installed, its 'from botocore.vendored import requests' resolves
against a module with no __path__ and every exception test in the worker
dies with "No module named 'botocore.vendored'" — ordering-dependent,
so green locally, red in CI workers.

Defenses (both, in depth):
- test_bedrock_adapter.py pre-imports the real botocore.exceptions at
  module scope, before any test can stub sys.modules — later imports are
  cache hits that can never re-execute the vendored import under a
  poisoned parent. Proven standalone: fake-parent repro fails without
  the pre-import, succeeds with it.
- autouse _boto_sys_modules_hygiene fixtures in all three files that
  plant fake boto* modules (adapter, integration, model-picker):
  snapshot every boto* sys.modules entry before each test, evict+restore
  after — no stub window can leak state into a later test regardless of
  worker ordering.
- importorskip targets botocore.exceptions (the module the tests
  actually need) instead of bare botocore, so a torn install skips
  instead of erroring.

148/148 across the four affected suites.
2026-08-22 19:30:10 -07:00
Teknium 0c435f4601 fix(update): reword refusal message — footgun linter matched prose 'venv open (' as bare open() 2026-08-22 19:30:10 -07:00
Teknium 83864c0b5d fix(update): a contended venv is never mutated — failed shim quarantine now refuses instead of warning (#87331)
The #87331 remaining half: when hermes.exe (or a sibling shim) could not
be renamed aside, the updater printed a warning and ran the installer
anyway — which died partway on the same locks and stranded the venv
between versions.

- _run_quarantined_install gains strict_quarantine: any shim whose
  rename failed every retry aborts BEFORE the install command runs
  (successful renames rolled back), raising ShimQuarantineError.
- The update dependency sync passes strict_quarantine=True. The update
  boundary turns the error into a refusal: defer via the
  update-incomplete marker, exit 2 (recorded as refused by the receipt
  net), never ZIP-fallback. Post-sync repair installs keep warn-and-try
  (their venv is already mutated; refusing buys nothing).
- The recovery installer (_install_repair._run_install_cmd) is strict
  unconditionally: marker survives, next launch retries after the
  holder exits.
- Live Windows E2E for the wine2e lane: a real child holds hermes.exe
  without FILE_SHARE_DELETE (the exact field lock shape), strict path
  refuses with zero installer invocations, releases roll back, and the
  same path proceeds once the holder exits.

Sabotage-verified: reverting the strict wiring makes both fail-closed
tests fail.
2026-08-22 19:30:10 -07:00
emozilla fe95ed3930 Merge origin/main: reconcile with PR #92092 (in-checkout launcher restore)
PR #92092 fixed the same vanished-launcher bug by restoring copies into
the legacy in-checkout hermes-agent\bin from the update tail. That
location is what this branch removes: untracked files there are swept
by the update autostash on every cycle (restore/sweep treadmill, plus a
parked stash entry per update under --keep-stash), and unconditional
exe copies break on relocatable venvs ('uv trampoline failed to
canonicalize script path'). This branch's managed-binary-dir layout
supersedes both mechanisms, so the merge resolves to it:

- drop _sync_windows_cli_launchers and its _ensure_acp_launcher call
  (Windows staging/repair lives in ensure_windows_bin_launchers at
  process start and migrate_windows_bin_path in the update tail);
  _ensure_acp_launcher is a Windows no-op again
- keep #92092's genuinely better installer semantics: staging stays in
  a dedicated Install-HermesCommandLaunchers function that throws
  BEFORE any PATH mutation when the required launcher cannot be staged
  and verified -- previously Set-PathVariable could put an empty dir on
  PATH and still print 'hermes command ready'. Reworked for this
  branch's layout: caller passes the destination ($HermesHome\bin),
  launcher form follows the venv (exe copy vs .cmd delegator), and the
  verify step accepts either form
- rework #92092's AST-lifted PowerShell test for the new function
  signature, keeping its fail-before-PATH-mutation assertions and
  adding relocatable-venv form-selection coverage
- drop tests/hermes_cli/test_windows_cli_launcher_repair.py (pinned the
  superseded in-checkout mechanism; equivalent and broader coverage
  lives in tests/hermes_cli/test_ensure_windows_bin_launchers.py)
2026-08-22 22:16:38 -04:00
Teknium 7a54ab22e6 fix(gateway): control-socket hardening from #92447 post-merge review
- bind under umask 0o177 so the socket is never world-connectable, even
  pre-chmod (review pt 3)
- verb handlers run in an executor: state-file reads stay off the
  adapter event loop (pt 2)
- inventory dedupes one multiplex gateway answering identify for several
  homes — one runtime record per pid, with regression test (pt 1)
- v1 wire contract (one request per connection) documented in the module
  docstring (pt 5); /tmp-unwritable skip in the short-home test

Live-verified: perms 600 at bind, identify 4.3ms via executor path, 20
rapid queries healthy.
2026-08-22 18:05:26 -07:00
Teknium 530028c213 test(gateway): pipe E2E compares the server's self-reported pid, tree-kills the uv trampoline
First windows-latest run proved the design premise in miniature: identify
answered pid 616 while Popen.pid said 8000 — uv's Windows python.exe is a
trampoline that spawns the real interpreter as a child, so the spawner's
PID view is wrong and the process's self-declaration is right. Assert
against the child's printed os.getpid(); use taskkill /T for teardown.
2026-08-22 16:45:00 -07:00
Teknium 67aac4d863 fix(gateway): control-socket fallback survives deep TMPDIR; tests bind-location-aware
CI runners put pytest tmp roots past sun_path, which routed every test
home through the fallback: two tests assumed in-home binding. Tests now
assert against the resolved bind location, and the fallback itself
prefers /tmp when tempfile.gettempdir() is too deep to fit sun_path.
2026-08-22 16:45:00 -07:00