Commit Graph

5936 Commits

Author SHA1 Message Date
Ishman82 b44c2bdab2 fix(desktop): spare concurrently starting backends
Desktop writes backend.lock.json only after a remote profile backend announces readiness. Concurrent Bot Mode profile starts could therefore see their young siblings as unowned PPID-1 processes and mutually reap them, causing SSH reconnect storms and stale turn leases.\n\nProtect unregistered backends for a bounded startup grace period, fail closed when age cannot be read, and cover young, unknown-age, boundary, lock-owned, and old-orphan behavior.
2026-08-23 00:52:50 -07:00
Teknium 8804e78354 feat(dashboard): Desktop and dashboard read the update receipt instead of inferring success (#91277 Phase-1 bullet 3)
Builds on @mrsucesso's durable-marker recovery (previous commit):

- GET /api/hermes/update/receipt — the full durable receipt (steps,
  skips, gateway restart outcome, fleet matrix) + compact summary; the
  authoritative update-outcome record (written by every run since
  #91283, including refused/failed).
- /api/actions/hermes-update/status now attaches the receipt summary,
  and when BOTH the in-memory registries and the update.log marker are
  gone (dashboard restarted + log rotated — the #81193 state), a
  finished receipt reports the outcome: success→0, partial→1. A
  still-running receipt proves nothing (clients keep polling).
- Desktop (updates.ts): the apply poll reads the attached receipt — a
  finished receipt whose run started at/after this apply is
  authoritative, replacing timeout-based failure inference across the
  update's restart gap ('Backend update failed' on successful updates,
  #81193; 'boot failed' during update restarts, #87359).

Live-verified: real uvicorn server + real UpdateReceipt writer (the
exact code hermes update runs) over real HTTP — receipt endpoint 200
with summary; #81193 state (no registries, no marker) reports success
from the receipt alone; partial receipt with a DOWN fleet row maps to
exit 1 (no false success).
2026-08-23 00:50:41 -07:00
Mauricio Ruiz 091106092b fix(dashboard): recover update success after restart
Persisted update completion markers survive the dashboard restart that clears in-memory action state. Recover the latest safe marker from update.log so remote Desktop clients do not report a successful backend update as failed.
2026-08-23 00:50:41 -07:00
Franci Penov e366df6889 fix(cli): treat a fork's upstream sync as an update
On a fork, `hermes update` compares HEAD against origin/main, and only then
syncs the fork from upstream — inside the `commit_count == 0` branch, which
returns immediately afterwards. So an update that pulls hundreds of commits
from upstream prints "Already up to date!" and skips everything the
post-update path does, including the dependency sync and the gateway restart.

Observed on a fork-based deployment: 1654 commits pulled, "Already up to
date!", and the launchd gateway left running. It then held pre-update modules
in memory while lazily importing post-update ones, and failed later with an
AttributeError for a method that plainly exists on disk — a mixed runtime that
looks nothing like an update problem. Correlating every run in update.log, a
restart happened on exactly the runs that pulled upstream *without* also
claiming to be up to date, and never once they started co-occurring.

Decide before the branch: capture HEAD, sync, and if HEAD moved, set
commit_count from the range so the normal post-update path runs. The pull that
follows is a no-op (the sync updates origin too); reaching the restart is the
point. commit_count is floored at 1 — HEAD moving *is* the update, so a failed
or zero count query must not send us back down the early return.

steps still being skipped afterwards.

Refs #73108
2026-08-23 00:19:46 -07:00
Teknium 1684877868 fix(update): a gateway killed by the restart phase and never replaced now fails the fleet check (DOWN row)
Phase-1 verification gap (#91277, found auditing our own landed matrix
against the mapped issues): collect_fleet_versions only listed gateways
with a LIVE pid, so 'restart stopped it and nothing came back' produced
NO row at all — the exact silent-failure shape the matrix exists to
catch (#88848/#74973 class) passed with exit 0.

- collect_fleet_versions(pre_restart_pids=...): a dead pid becomes a
  'down' row only when it was alive at update start AND its runtime
  status still claims a running state. Rollout-safe: no snapshot (old
  callers), clean stops, startup failures, and stale records from
  long-dead gateways keep the historical no-row behavior.
- print_fleet_version_matrix escalates on down rows like stale ones
  (exit 1) with the per-profile restart remediation.
- cmd_update passes its existing pre-restart PID snapshot.

Sabotage-verified (reverting the membership check fails the new test);
live-verified with a real spawned-then-killed process producing the
DOWN row and matrix escalation.
2026-08-22 23:46:06 -07:00
Josh Holt 5f0a8f8739 fix(picker): harden keyless provider gate logging and credentials path validation 2026-08-22 23:25:43 -07:00
Josh Holt b9f17ba3f1 fix(picker): scope Vertex explicit-config to Hermes signals, not ambient ADC
Address review feedback: the gate reused has_vertex_credentials(), which also
returns True for an ambient GOOGLE_APPLICATION_CREDENTIALS path. That var is
commonly set globally for unrelated GCP work, so a user who never configured
Hermes for Vertex would see it in the explicit-only picker and could spend
against those credentials — weakening the explicit-configuration guarantee the
gate documents (mirrors the existing _IMPLICIT_ENV_VARS carve-out).

Add has_explicit_vertex_config() in agent/vertex_adapter.py that checks only
Hermes-scoped signals — VERTEX_PROJECT_ID / vertex.project_id (project
override) or a resolvable VERTEX_CREDENTIALS_PATH — and NOT
GOOGLE_APPLICATION_CREDENTIALS. Route the auth gate through it.

Adds a regression test asserting an ambient GOOGLE_APPLICATION_CREDENTIALS
path alone does not mark Vertex explicit, and updates the existing test to
drive the real config signal instead of mocking has_vertex_credentials().
2026-08-22 23:25:43 -07:00
Ajad van Wyk 3503c06d80 fix: surface Bedrock in explicit-only model pickers when AWS env credentials are set
is_provider_explicitly_configured() only checked provider env vars for
auth_type="api_key" providers. Bedrock is registered with
auth_type="aws_sdk" and an empty api_key_env_vars tuple, so a user who
sets AWS_BEARER_TOKEN_BEDROCK (or an AWS_ACCESS_KEY_ID +
AWS_SECRET_ACCESS_KEY pair) in .env was never counted as having
explicitly configured the provider.

Symptom: the desktop model picker (and any consumer of
build_models_payload(explicit_only=True)) silently hid the Bedrock row
even though list_authenticated_providers had discovered credentials and
built a full model list for it. Reproduced on main:

    build_models_payload(ctx, explicit_only=False)
      -> ['moa', 'nous', 'bedrock', ...]        # row exists, 132 models
    build_models_payload(ctx, explicit_only=True)
      -> ['nous', ...]                          # bedrock filtered out

Fix: aws_sdk-type providers now count as explicitly configured when
Bedrock-relevant env credentials are present. Deliberately env-var-only:
ambient sources (AWS_PROFILE / SSO, EC2 IMDS, container credentials)
still do NOT auto-surface, consistent with the gate's purpose (#56974)
and with a lone AWS_ACCESS_KEY_ID (no secret) not counting.

Tests: six behavior-contract cases in test_auth_provider_gate.py
covering bearer token, key pair, lone key id, ambient AWS_PROFILE,
no-credential baseline, and non-leakage into other providers.
2026-08-22 23:25:43 -07:00
Kevin Yin 9ce46a0968 fix(models): bind Anthropic pool key to its endpoint 2026-08-22 23:25:29 -07:00
Kevin Yin 8ede2e1472 fix(models): discover Anthropic pool API keys 2026-08-22 23:25:29 -07:00
Teknium 4f0e466e5d fix: accept pool-only Anthropic OAuth entries in the desktop picker filter
The salvaged carve-out covered the PKCE file and Claude Code credentials
but missed the canonical wired-token location — auth.json
credential_pool.anthropic oauth entries. Discovery accepts those via
pool.has_credentials(), so the filter must too. Read-only dict access;
api_key pool entries intentionally stay excluded (E2E case 4).
2026-08-22 22:00:19 -07:00
Viktor 4ec57d56a9 fix(inventory): keep Anthropic OAuth logins visible in desktop pickers
The desktop explicit-only picker filter drops any provider row where
is_provider_explicitly_configured() is False. That gate only recognizes
active_provider, model.provider, and API-key env vars — so a user who
authenticated Anthropic via OAuth (Hermes device flow or Claude Code
~/.claude/.credentials.json) had the Anthropic row silently hidden from
the desktop model picker even though list_authenticated_providers()
accepted those same credentials when building it (model_switch.py has an
equivalent special-case in its discovery path).

Unlike ambient CLI tokens (gh -> copilot), an OAuth access token only
exists after an interactive login, so its presence is deliberate user
configuration. Add a narrow carve-out for the anthropic slug that keeps
the row when either OAuth reader finds an accessToken; the strict gate
still governs every other provider, and is_provider_explicitly_configured
itself stays untouched so PR #4210's consent guard for auxiliary tasks
keeps its current behavior.
2026-08-22 22:00:19 -07:00
Teknium 8b86097a62 fix(gateway): honor env-configured local backends and persist Tool Gateway declines
Follow-ups on top of the salvaged #92665 commits, closing the two gaps
called out in #92647:

- SEARXNG_URL / CAMOFOX_URL now count as direct web/browser configuration
  in _get_gateway_direct_credentials(), so env-only keyless local setups
  are offered unchecked instead of pre-checked.
- Submitting the checklist with unconfigured tools left unchecked records
  them in tool_gateway_declined_tools (new known root key) and stops
  pre-checking them on subsequent Nous model swaps; opting in later clears
  the decline. Cancel (Ctrl-C/ESC) records nothing.
- has_direct labels mention SearXNG/Camofox when that's what's detected.
- Docs: tool-gateway.md gains an enablement-checklist section.
2026-08-22 21:36:55 -07:00
chelsealong ce9ddd35e7 fix(gateway): fix 3-tuple/4-tuple arity crash in get_gateway_eligible_tools
The early "fail closed" returns (account fetch error, not entitled,
non-nous provider) still returned 3-element tuples after the function's
happy path and every caller moved to a 4-tuple, crashing
prompt_enable_tool_gateway with ValueError for any logged-in Nous
account that isn't paid or pool-entitled — the common case, hit
unconditionally from `hermes model`.
2026-08-22 21:36:55 -07:00
chelsealong 4edb24276f fix(tools): don't pre-check keyless local backends in Tool Gateway checklist
get_gateway_eligible_tools() classified a tool as "unconfigured" whenever
it found no direct API-key credential, ignoring an explicit non-nous
selection already stored in config (e.g. web.backend: searxng, a keyless
self-hosted backend). Every unconfigured tool is pre-checked in the
`hermes model` Tool Gateway checklist, so a single Enter during Nous model
setup silently rewrote web.backend (and similarly browser.cloud_provider,
tts/stt/image_gen.provider) to "nous" for users who had deliberately
configured a keyless local backend.

get_gateway_eligible_tools() now also resolves each tool's stored
selection via the existing _selected_provider() helper and routes an
explicit non-nous selection into a new explicit_configured bucket instead
of unconfigured. prompt_enable_tool_gateway() never offers those tools,
so they can no longer be pre-checked or accidentally overwritten.

Fixes #92647
2026-08-22 21:36:55 -07:00
Teknium f4067774aa fix(update): token-based control-plane classifier + live E2E for the Desktop-lifecycle cold-start skip (#76129 salvage follow-up)
On top of @686f6c61's premise-corrected #76745:

- _looks_like_desktop_control_plane now uses the parser-derived
  _hermes_holder_subcommand instead of substring matching — the
  #90778/#91869 class ('-m dashboard chat' and 'kanban --preserve-cache'
  argv no longer read as control planes). Regression test added,
  sabotage-verified (reverting to substrings fails it).
- Live E2E (this host, real processes + real spawn ledger): live
  supervised serve owns lifecycle; killed spawner (orphan) does not;
  dead serve entry excluded; empty ledger does not.
- Live Windows E2E for the wine2e lane: real self-registered ledger
  entry suppresses the actual cold-start plan; dead serve restores it;
  holder-scan fallback rung proves the token classifier live.

Co-authored-by: 686f6c61 <github@00b.tech>
2026-08-22 21:36:44 -07:00
686f6c61 4ccc4b6931 fix(update): skip Windows gateway cold-start when Desktop owns lifecycle
Vestigial autostart is not proof the user wants a standalone gateway
run. When Desktop currently supervises this install's control plane,
the updater must not spawn a competing messaging daemon. Serve is not
treated as gateway-equivalent.
2026-08-22 21:36:44 -07:00
Teknium c9c44d0df9 Merge pull request #92636 from NousResearch/fix/windows-launcher-managed-bin
fix(windows): stage hermes launchers in the managed binary dir, not the git checkout
2026-08-22 20:58:00 -07:00
emozilla 0b01599eef fix(tests): repair the two Linux-lane CI failures on this branch
Both failed only in the full Linux suite, which the targeted local
battery never ran:

- test_update_zip_two_phase.py's AST guard (#76105) flags any code
  literal "Scripts" in hermes_cli as an open-coded venv layout.
  migrate_windows_bin_path's legacy PATH key now derives it via
  venv_bin_dir(root / "venv", windows=True) — same value, canonical
  helper. The literal `venv` component stays: the key must match what
  the pre-#83797 installer wrote to the registry, not where the venv
  lives now.
- The managed-bin marker tests built expected PATH entries from
  tmp_path, so on a POSIX host they compared forward-slash strings
  against the backslash markers and could never match. Markers match
  Windows registry PATH entries, so the tests now feed Windows-shaped
  literals — host-independent, same contract.
2026-08-22 23:10:32 -04:00
Teknium 0c435f4601 fix(update): reword refusal message — footgun linter matched prose 'venv open (' as bare open() 2026-08-22 19:30:10 -07:00
Teknium 83864c0b5d fix(update): a contended venv is never mutated — failed shim quarantine now refuses instead of warning (#87331)
The #87331 remaining half: when hermes.exe (or a sibling shim) could not
be renamed aside, the updater printed a warning and ran the installer
anyway — which died partway on the same locks and stranded the venv
between versions.

- _run_quarantined_install gains strict_quarantine: any shim whose
  rename failed every retry aborts BEFORE the install command runs
  (successful renames rolled back), raising ShimQuarantineError.
- The update dependency sync passes strict_quarantine=True. The update
  boundary turns the error into a refusal: defer via the
  update-incomplete marker, exit 2 (recorded as refused by the receipt
  net), never ZIP-fallback. Post-sync repair installs keep warn-and-try
  (their venv is already mutated; refusing buys nothing).
- The recovery installer (_install_repair._run_install_cmd) is strict
  unconditionally: marker survives, next launch retries after the
  holder exits.
- Live Windows E2E for the wine2e lane: a real child holds hermes.exe
  without FILE_SHARE_DELETE (the exact field lock shape), strict path
  refuses with zero installer invocations, releases roll back, and the
  same path proceeds once the holder exits.

Sabotage-verified: reverting the strict wiring makes both fail-closed
tests fail.
2026-08-22 19:30:10 -07:00
emozilla fe95ed3930 Merge origin/main: reconcile with PR #92092 (in-checkout launcher restore)
PR #92092 fixed the same vanished-launcher bug by restoring copies into
the legacy in-checkout hermes-agent\bin from the update tail. That
location is what this branch removes: untracked files there are swept
by the update autostash on every cycle (restore/sweep treadmill, plus a
parked stash entry per update under --keep-stash), and unconditional
exe copies break on relocatable venvs ('uv trampoline failed to
canonicalize script path'). This branch's managed-binary-dir layout
supersedes both mechanisms, so the merge resolves to it:

- drop _sync_windows_cli_launchers and its _ensure_acp_launcher call
  (Windows staging/repair lives in ensure_windows_bin_launchers at
  process start and migrate_windows_bin_path in the update tail);
  _ensure_acp_launcher is a Windows no-op again
- keep #92092's genuinely better installer semantics: staging stays in
  a dedicated Install-HermesCommandLaunchers function that throws
  BEFORE any PATH mutation when the required launcher cannot be staged
  and verified -- previously Set-PathVariable could put an empty dir on
  PATH and still print 'hermes command ready'. Reworked for this
  branch's layout: caller passes the destination ($HermesHome\bin),
  launcher form follows the venv (exe copy vs .cmd delegator), and the
  verify step accepts either form
- rework #92092's AST-lifted PowerShell test for the new function
  signature, keeping its fail-before-PATH-mutation assertions and
  adding relocatable-venv form-selection coverage
- drop tests/hermes_cli/test_windows_cli_launcher_repair.py (pinned the
  superseded in-checkout mechanism; equivalent and broader coverage
  lives in tests/hermes_cli/test_ensure_windows_bin_launchers.py)
2026-08-22 22:16:38 -04:00
Teknium 7a54ab22e6 fix(gateway): control-socket hardening from #92447 post-merge review
- bind under umask 0o177 so the socket is never world-connectable, even
  pre-chmod (review pt 3)
- verb handlers run in an executor: state-file reads stay off the
  adapter event loop (pt 2)
- inventory dedupes one multiplex gateway answering identify for several
  homes — one runtime record per pid, with regression test (pt 1)
- v1 wire contract (one request per connection) documented in the module
  docstring (pt 5); /tmp-unwritable skip in the short-home test

Live-verified: perms 600 at bind, identify 4.3ms via executor path, 20
rapid queries healthy.
2026-08-22 18:05:26 -07:00
Teknium 60b6269142 feat(gateway): gateway-owned control socket — identify/status verbs, fleet consumers prefer it over scans (#92091 step 1)
The gateway now creates a local control socket at startup (Unix domain
socket at $HERMES_HOME/gateway.sock with a pointer-file fallback for
long paths; named pipe on Windows) and answers versioned JSON verbs:

- identify: pid, profile, hermes_home, code_sha/code_version (#91283
  stamps, now queryable live), self-declared supervisor kind, start_time
- status: the live runtime-status payload, answered by the process itself

Bound immediately after the PID-file O_EXCL claim (the moment the
process becomes the authoritative gateway for its HERMES_HOME), removed
on clean shutdown; a successor clears any stale socket on bind. Strictly
non-fatal: bind failure only means consumers use the old path.

Consumers migrated (observability only, scan layer demoted to fallback,
never deleted):
- collect_fleet_versions() (post-update fleet matrix): prefers a live
  identify answer over gateway_state.json; entries carry source=socket
- collect_runtime_inventory() (hermes update --plan): prefers the
  socket, and takes the gateway's own supervisor declaration instead of
  inferring it from PID scans

Old gateways mid-upgrade, crashed processes, and bind failures behave
exactly as before. Never a TCP port; filesystem/pipe ACLs are the auth
boundary (0600 socket).

Part of #91277 (fleet-update reliability). Design: #92091.
2026-08-22 16:45:00 -07:00
Bruce Xu 50bbcbf2b4 fix(state): fail closed on unscoped SQLite corruption 2026-08-23 02:55:03 +05:30
emozilla 5d5179d7a9 fix(windows): remove the hermes launchers on uninstall
Every uninstall mode deletes the code checkout, but the launchers in
the managed binary dir (%LOCALAPPDATA%\hermes\bin) live outside it and
survived -- so `hermes` in a new terminal resolved to a launcher whose
venv target was gone and errored, which reads worse than
command-not-found.

remove_windows_bin_launchers deletes both launcher forms (.exe/.cmd)
from the managed binary dir in every uninstall mode, anchored on the
default Hermes root so profile sessions cannot redirect the sweep into
profiles\<name>\bin. When the uninstall itself runs through the
launcher, that exe is mandatory-locked against deletion but not rename
(the same fact _quarantine_running_hermes_exe relies on), so it falls
back to renaming the launcher aside.

The managed uv (uv*.exe) in the same dir survives, and the hermes\bin
PATH entry is swept only on a full wipe from the default root
(include_managed_bin) -- a keep-data uninstall keeps the still-working
uv resolvable for reinstalls.

A lockstep test parses install.ps1's staging loop so the swept names
cannot drift from the staged names silently.
2026-08-22 13:39:04 -04:00
emozilla 679e9cd294 fix(windows): stage hermes launchers in the managed binary dir, not the git checkout
The installer staged the hermes/hermes-acp launcher copies at
hermes-agent\bin -- inside the git working tree -- and put that dir on
the user PATH (#84452). The update command's pre-pull autostash
(git stash push --include-untracked) swept those untracked, unignored
copies off disk, and once the desktop updater stopped re-applying
stashes (--keep-stash, 5dd221d442) nothing restored them: `hermes`
stopped resolving in every new terminal on every desktop-updated
install.

Move the canonical launcher home to the managed binary dir
(%LOCALAPPDATA%\hermes\bin, next to the managed uv) -- outside the
checkout, where no git operation can ever touch it. The dir is
per-machine and shared by every profile, so all anchoring uses
get_default_hermes_root(), never HERMES_HOME (which points inside
profiles\<name> under `hermes -p`).

The copy design also had a second latent break: managed-uv rebuilds
create relocatable venvs, and a relocatable venv's exe trampoline
resolves relative to its own location -- a copy outside venv\Scripts
dies with 'uv trampoline failed to canonicalize script path'. Launcher
form now depends on the venv (lockstep in install.ps1 and
_install_repair.py): exe copy for normal venvs, a .cmd delegator
invoking the in-venv exe by absolute path for relocatable ones. Either
form counts as present, so pre-rebuild exe copies are left alone.

Delivery to the existing fleet, per cohort:

- already-broken installs cannot run the CLI, so an import-time heal in
  hermes_cli.main (ensure_windows_bin_launchers) re-stages missing
  launchers when the desktop app spawns its backend -- the one channel
  that still reaches them. Gates fail toward inaction: canonical dir
  only for the managed clone, legacy hermes-agent\bin only while the
  user PATH still resolves through it (some pre-managed-uv installs
  have no hermes\bin PATH entry; the legacy re-stage is what fixes
  those). Staging-name + os.replace keeps concurrent process starts
  from tearing a launcher; the helper never raises.
- healthy old-layout installs migrate in the update tail
  (migrate_windows_bin_path): stage canonical launchers, verify them
  BEFORE touching the registry, prepend hermes\bin to the user PATH,
  strip the legacy entries (hermes-agent\bin and venv\Scripts, #83797),
  preserving REG_EXPAND_SZ and raw %VARS%. The legacy dir's files stay
  on purpose -- configs holding absolute launcher paths keep working;
  only the sweepable PATH resolution route goes.
- fresh installs get the new layout from install.ps1 directly.

/bin/ is gitignored so the one update that DELIVERS this fix cannot
sweep pre-migration launchers a final time under the old rules; the
gitignore line, the legacy re-stage branch, and the update-tail call
are transition machinery with a named expiry once the fleet has
migrated.

Also rewrites _ensure_acp_launcher's stale Windows paragraph to match
(raw docstring fixes its invalid \S escape) and updates the Windows
native docs to the new layout, with a docs<->installer parity test.
2026-08-22 13:38:56 -04:00
poisdahl a36d6704b6 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821 2026-08-22 17:31:04 +02:00
poisdahl fd41164861 fix(history): keep carrier rewinds race-safe after refresh 2026-08-22 17:30:35 +02:00
Kshitij Kapoor 3616145723 fix(gateway): keep loop-watchdog default at 3 strikes; dedupe constants; register knobs in config defaults
Downscope of the salvaged #89134 per review: the 3->8 default raise was
symptom tolerance for the false-positive class the off-loop heartbeat +
two-witness probe fixes at the root — fleet-wide it would only delay
genuine-wedge recovery ~2.7x. The three tuning knobs keep independent
operator value and stay:

- default max_strikes back to 3 everywhere (constant, dataclass,
  from_dict fallback, floor clamp, tests)
- gateway/config.py + gateway/run.py now reference the
  shutdown_watchdog DEFAULT_* constants instead of duplicating literals
  in three places (drift hazard)
- knobs registered in hermes_cli/config_defaults.py alongside the
  sibling gateway.loop_watchdog bool
2026-08-22 20:40:23 +05:30
poisdahl a5b326a471 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821
# Conflicts:
#	tests/agent/test_reference_handoff_active_turn.py
2026-08-22 16:47:39 +02:00
Kshitij Kapoor a18f90847b refactor(gateway): promote parse_systemd_duration_to_us to public
hermes_cli/gateway.py's restart-wait sizing (from #92175) was the only
cross-module import of an underscore-private shutdown_forensics helper.
Promote it (private alias retained for existing patchers).
2026-08-22 16:26:51 +05:30
fangliquan a4f16e3fef fix(gateway): retain failed replacement evidence 2026-08-22 15:27:26 +05:30
fangliquan 596bfc557f fix(gateway): distinguish failed systemd replacements 2026-08-22 15:27:26 +05:30
fangliquan 83b09ebd0a fix(gateway): make handoff recovery idempotent 2026-08-22 15:27:26 +05:30
fangliquan 91fb175188 fix(gateway): preserve systemd handoff recovery 2026-08-22 15:27:26 +05:30
fangliquanflq 5b024c7ccc fix(gateway): make systemd the sole restart owner 2026-08-22 15:27:26 +05:30
Gille a08b909199 fix(windows): preserve launcher layout invariants 2026-08-22 00:05:51 -07:00
Gille 9782275b2a fix(windows): restore dedicated CLI launchers on update 2026-08-22 00:05:51 -07:00
Teknium 8f30e9c77a feat(desktop): Send Diagnostics — one-click redacted debug-bundle upload from the error card
New diagnostics.share_nous RPC reuses the CLI --nous pipeline
(collect_share_bundle → build_nous_bundle → share_to_nous) with redaction
forced on; accepts redacted error context + client-side extra files
(local desktop.log on remote connections) with sanitized labels and size
caps. Desktop: Send Diagnostics action on the failed-turn error card →
consent modal (privacy notice, explicit Upload) → private view link +
GitHub Issues / Nous Portal Support / Discord handoff. CLI --nous success
output gets the same three-destination pointer. i18n en/ja/zh/zh-hant/ar;
docs updated.
2026-08-21 23:01:30 -07:00
JonthanaHanh 01c14ad7f3 fix(update): ZIP swap preserves the built desktop app (apps/desktop/release)
The #70337/#87331 win-unpacked wipe half, from PR #70477 by @JonthanaHanh
(reimplemented against the two-phase staged swap that postdates that
branch — the live release/ dir is grafted into the staged apps copy
BEFORE the atomic commit, so preservation rides the same rollback
machinery instead of a post-hoc copy).

Co-authored-by: JonthanaHanh <92574114+JonthanaHanh@users.noreply.github.com>
2026-08-21 22:08:16 -07:00
kshitijk4poor eac3f645ef fix(update): don't ZIP-fallback on dependency failures or dirty trees
Surgical reapply of PR #87878 (@kshitijk4poor's salvage of #87327 by
@liruixinch) onto current main — the receipt-boundary and summary
changes from this session made the original commits conflict.

- ZIP fallback now keys on git ACTUALLY having failed
  (_should_zip_fallback_on_update_error): a dependency-install failure
  after a successful pull can't be fixed by re-downloading source and
  would clobber the tree (#87331 cascade trigger, #87304).
- _abort_zip_update_if_dirty_tree: refuse to overlay a dirty checkout
  (-uall so user gitconfig can't blind the guard) + pre-swap TOCTOU
  re-check with our own staging artifacts filtered (#91962, #87304).
- Failure-stage naming (_format_update_failure_stage) + stderr tail so
  'Git update failed' stops mislabeling pip/uv failures.
- Receipt finalize preserved on the no-fallback failure path.

Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
Co-authored-by: liruixinch <liruixinch@outlook.com>
2026-08-21 22:08:16 -07:00
Teknium 7d6db4efb8 fix(update): holder classifier derives value-flags from the real parser; de-flake goal-resume fixture
Review on #91869 (@andrexibiza): the handwritten value_flags subset
misparsed '--reasoning high serve' as subcommand 'high' and
'-m dashboard serve' as 'dashboard' — recreating the wrong-hint class.
_holder_value_flags() now introspects build_top_level_parser() (every
option with nargs != 0, plus the pre-argparse profile selectors), with
a static fallback for broken-tree updates, --flag=value handled.
Regressions for --reasoning/-m/-t/--model=/-c per review.

De-flake test_goal_resume_restart: the fixture only set the HERMES_HOME
env var, but get_hermes_home() prefers the context-local override — an
override leaked by any earlier test in the xdist worker pointed the
goals DB at a dead tmp dir and resume enqueued nothing (the CI-only
red). Fixture now pins the override via set/reset_hermes_home_override.
Mechanism proven both ways: env-only fixture cannot beat a leaked
override; pinned fixture immune.
2026-08-21 19:11:55 -07:00
Teknium c02cac00ce fix(update): venv-holder labels parse the real subcommand; gateway ancestors stay visible to the scan
#90778: _hermes_holder_subcommand() — token-based parse of the actual
Hermes subcommand (profile selectors skipped, flags never matched), so
'hermes dashboard' stops being labeled as the Desktop backend and
'--preserve-cache' stops matching 'serve'. Unknown argv gets no hint
instead of a wrong one.

#87594: ancestor-exclusion in _detect_venv_python_processes and
_venv_launcher_ancestors now carves out GATEWAY ancestors (canonical
looks_like_gateway_command_line): when /update runs as the gateway's
child, the gateway stays visible to the scan so the pause machinery can
stop it, while shells/terminals/own-venv ancestry stay excluded.

15 cross-platform classifier tests; live Windows E2E suite is the
acceptance gate on this branch.
2026-08-21 19:11:55 -07:00
AlexFucuson9 f5adeed3dc fix(cli): guard empty message text in _display_resumed_history
text.splitlines() returns [] for empty strings. Accessing msg_lines[0]
then raises IndexError, making session resume crash when the session
contains a message with empty or whitespace-only text (e.g. reasoning-only
turns, tool-only assistant messages).

Guard with `or [""]` in all three branches (user, assistant_last,
regular assistant) so an empty message renders as a blank line.

Fixes #59265

Co-authored-by: AlexFucuson9 <AlexFucuson9@users.noreply.github.com>
2026-08-22 03:59:28 +05:30
Teknium 04acfb9673 fix: remove function-level 'import time as _time' that shadowed the module import
The in-function import made _time local to all of _cmd_update_impl, so
the orphan-backend reap path (which runs earlier in the function) hit
UnboundLocalError before the import line executed. The module-level
'import time as _time' at the top of update_cmd.py already covers the
divergence-merge safety tag.
2026-08-21 15:23:49 -07:00
Teknium 70151dd549 feat(update): updates.parked_branch_strategy gates the in-place merge; switch stays the default
Adapts the in-place branch update from PR #89507 (@willfrombr) onto the
switch-by-default behavior: the deterministic switch path remains the
default so non-interactive updates (desktop, gateway, cron) never dead-end
on a merge conflict, and deliberate custom-branch users opt in with
updates.parked_branch_strategy: update_in_place. --switch-branch overrides
the in-place strategy for one run (deep feature branches that must not
accumulate update merge commits). Docs + config comments + tests cover
all three routes.

Co-authored-by: Willian Santos <285090322+willfrombr@users.noreply.github.com>
2026-08-21 15:23:49 -07:00
Willian Santos 4fad27a101 feat(update): --switch-branch opts an unmerged branch out of the in-place merge
Review feedback on #89507: in-place merging suits a branch that tracks the
target with a small patch set, but a long-lived feature branch (a PR branch
hundreds of commits deep) does not want an update-driven merge commit
written into its history. Reported against a checkout carrying 819 unmerged
commits.

--switch-branch routes the unmerged case to the switch path instead: the
checkout moves to the update target and updates there, and the branch is
left byte-identical — no merge, no commit, nothing written to it. The tree
is known clean on that path (the guard checks dirty before cherry), so a
dirty tree still gets the loud skip, unchanged.

Opt-in: without the flag the default remains the in-place update, which is
what keeps a small-patch-set branch's running code current.

Tests: the flag switches and leaves the branch tip byte-identical; the
default without it still updates in place. The first fails if the flag's
branch is severed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 15:23:49 -07:00
Willian Fernandes Santos 91096bb2f0 feat(update): update branches carrying unmerged commits in place instead of skipping
The parked-branch guard (8ce8ffd429) distinguishes checkouts by what the
branch carries, then treats both non-clean cases the same: a stale
fully-merged leftover is switched back to the target (correct), but a
branch with unmerged commits — a branch someone is actually working on —
gets CODE UPDATE SKIPPED and exit 1. For anyone running a maintained
custom branch on top of main, every update now refuses, and the guidance
('checkout main') abandons their branch.

The guard's own reason codes already separate the cases, so use them:

- fully merged      -> switch back to the target (unchanged)
- unmerged:N        -> update the branch IN PLACE: fetch, then bring
                       origin/<target> into the checkout. Fast-forward
                       when possible; on divergence, a true merge behind
                       a pre-update safety tag, stopping cleanly on
                       conflict. The checkout never moves; local commits
                       survive; the running code advances.
- dirty/unverifiable/opted out -> skip loudly (unchanged)

The post-pull success gate learns that an in-place update legitimately
ends on a non-target branch: origin/<target> was merged INTO the checkout,
so refusing to claim success there would fail every update that did
exactly the right thing.

Guard tests updated: the unmerged case now asserts the in-place outcome —
target code arrives (b.txt from c3), the branch's own commit survives, and
HEAD never moves. 18/18 guard tests, 20/20 with the diverged-update suite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 15:23:49 -07:00
Teknium bbbc50acc2 fix: hermes update no longer strands non-interactive updates on a parked branch with unmerged commits
A clean checkout parked on a feature branch now always switches to the
update target. Unmerged commits are safe on the branch (git checkout
never discards committed work) and get a loud 'kept' notice naming the
branch, count, and the checkout command to resume the work. Previously
the update hard-skipped with exit 1 — a dead end for the desktop update
button, gateway /update, and cron, which have no way to resolve a skip.

Dirty trees (uncommitted changes) still skip loudly, and the
updates.auto_switch_parked_branch: false opt-out still pins the branch.
2026-08-21 15:23:49 -07:00