Commit Graph

5936 Commits

Author SHA1 Message Date
nftpoetrist 4d4cffd118 fix(profiles): make_targz writes to a temp file and renames, not the destination directly
tarfile.open(archive_path, "w:gz") truncates the destination the instant
it opens. If tf.add() fails partway (disk full, permission loss,
interruption), whatever was previously at that path is gone — including
an existing profile or board export the caller chose to overwrite. This
is the same failure shape a7e7de6407 just fixed for the desktop gateway
file-save path, one commit earlier in the same window, but it was never
propagated to this shared archive-writing primitive even though board
export gained a new caller into it in that same window.

make_targz now writes into a sibling temp file (mkstemp, same directory
as the destination so the final step is a same-volume rename) and only
replaces the destination via os.replace() after the archive is fully
written and closed, mirroring the mkstemp+os.replace pattern already
used throughout this codebase (agent/secret_sources/_cache.py,
cron/jobs.py, gateway/status.py, etc). The temp file is unlinked on any
failure.
2026-08-31 11:19:05 -07:00
RickyYii 3145986c20 fix(cli): honour model_aliases api_key, stop cross-provider key leak (#83612)
Salvaged from PR #84199 by @RickyYii. DirectAlias gains api_key/key_env; the direct-alias override re-resolves credentials against the alias endpoint (host-gated, #28660) and reuses the pre-alias key only on an origin match; oneshot -m <alias> passes the alias key as explicit_api_key; direct-alias branch gains the OLLAMA_API_KEY host gate. Fixes #83612.
2026-08-31 10:59:45 -07:00
Teknium ce942ab125 fix(update): fingerprint orphan backends from the classification psutil handle
The orphan-backend classifier fingerprinted candidates via
gateway.status.get_process_start_time, which prefers /proc/<pid>/stat —
the HOST process table, in clock ticks. Under the fake-psutil test harness
(and any containerized run where the PID number happens to exist on the
host) that returns the WRONG process's fingerprint in the WRONG units,
while pid_is_hermes verifies via psutil centiseconds at kill time: the
guard would then refuse every legitimate reap. Read create_time() from the
same psutil handle used for classification, quantized exactly like
gateway.status does on Windows, so the fingerprint round-trips.

Also covers the Windows-lane sibling: test_uses_netstat_and_taskkill_on_windows
now pins the guarded call path, plus a new refusal test for a non-bridge
listener PID (#89614 class).
2026-08-31 10:41:54 -07:00
Teknium 90e916efc9 fix(windows): compose the taskkill identity guards into one fail-closed class fix
Salvage hardening on top of the three cherry-picked contributor commits
(#91297 gebilaowang404 + AlexMnrs, #96741 burak33bb, #98826 ayushnangia),
closing the remaining unverified-PID kill sites as one class (#98814, #89614):

- pid_is_hermes: token-boundary 'hermes' match (no more loose substring
  false-positives), and an explicit start-time expectation is now honored
  on POSIX too (a mismatched fingerprint is a recycled PID on any platform).
- kill_process_tree: drop the guard on our OWN retained Popen child — a
  retained handle pins the PID, so the check could only false-refuse.
- gateway.status.terminate_pid: POSIX force-kills also refuse when a
  caller-provided expected_start_time no longer matches.
- kill_gateway_processes: re-verify the LIVE cmdline at kill time (the
  scan-time match is a TOCTOU window).
- _reap_unsupervised_gateway_orphans: fingerprint orphans at scan time and
  require a still-matching identity before the delayed SIGKILL escalation.
- whatsapp _kill_port_process: never kill a bare netstat/lsof-scanned PID
  unless the live process is actually a node bridge (was a stranger-kill).
- browser daemon reap/close paths: pass the start-time fingerprint into
  ProcessRegistry._terminate_host_pid (previously unverified), and the
  session-close path now runs the same daemon identity verification as
  the orphan reaper.
- tests/hermes_cli/test_taskkill_identity_windows_live.py: live Windows
  probes (real spawned processes, real psutil ancestry) wired into the
  on-demand windows-latest wine2e lane.

Fixes #98814
Fixes #89614
2026-08-31 10:41:54 -07:00
Ayush Nangia c923b53913 fix(update): refuse gateway ancestor tree-kill on Windows 2026-08-31 10:41:54 -07:00
burak33bb ed6d5fc803 fix(windows): require process identity before taskkill 2026-08-31 10:41:54 -07:00
gebilaowang404 cdd063528f fix(hermes_cli): fail-closed PID-ownership guard before Windows taskkill
Guard every Windows `taskkill /PID` against stale/recycled PIDs
(#89614: 8x 0xEF blue screens; a rebooted PID can be svchost.exe).

Adopted the community patch by AlexMnrs (commit 0162465): shared
psutil-based (pid, create_time) guard reusing the repo's existing
get_process_start_time machinery:
- fail closed on invalid/unknown/recycled identities (0/-1/None/bool/non-int)
- capture identity at discovery, re-validate at kill time
- all three sites through pid_is_hermes; taskkill stays hidden

Sites: _subprocess_compat.kill_process_tree,
dashboard_procs._kill_stale_dashboard_processes (win32),
update_cmd._stop_process_trees.

Refs #90471, #89614

Co-authored-by: Alex Monrás <AlexMnrs@users.noreply.github.com>
2026-08-31 10:41:54 -07:00
liuhao1024 db2fd5f59a fix(cli): answer clarify headless in single-query turns
hermes chat -q wired the interactive prompt_toolkit clarify callback
unconditionally, but a -q turn never builds the prompt_toolkit
application — the modal can never be painted or answered, so the turn
polls its response queue until agent.clarify_timeout expires (default
3600 s, 0 = unlimited). The gateway, cron jobs, the kanban dispatcher
and inter-agent wakeups all deliver work as -q turns. Route the
single-query case to a headless callback at the agent-construction site
that already knows _single_query_mode, mirroring _oneshot_clarify_callback
on the -z path (#94943; third member of the family after #86909 and
#88013).
2026-08-31 10:09:42 -07:00
webtecnica 839de43d52 fix(cli): stop raw CSI bytes from Shift+Space leaking into buffer (#88071) 2026-08-31 10:08:48 -07:00
phi hu 081030a7dd fix(config): warn for empty platform toolsets 2026-08-31 10:08:31 -07:00
phihu ccd32a9f0b fix(config): warn when a platform_toolsets entry is an empty list
validate_platform_toolsets() accumulated a single valid_count across every
platform, so the "zero valid toolsets" safety net was suppressed as soon as any
one platform carried a valid toolset. A platform wiped to [] — the active one,
typically cli — therefore produced no warning at all.

resolve_enabled_toolsets() honours that empty list verbatim ([] is a list, so
the platform-default fallback is skipped), leaving the agent with zero tool
schemas. The model then has nothing to call and emits the tool call as
assistant text with finish_reason=stop: no error, no warning, no log entry.
That is the silent-failure mode this module was written to prevent (#38798).

Note the asymmetry this leaves intact: a malformed *string* value is not a list,
so it falls back to the platform default and fails open (#78103); an empty list
fails closed. The fail-closed resolution is deliberate (the explicit_empty_
selection contract in tools_config.py, and #82010 wants it persistable), so this
only adds the missing warning and does not change resolution semantics.

Fixes #89050

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 10:08:31 -07:00
zengzheqing 98e2f110cf fix(curator): restore complete skill packages on ledger rollback (#96962)
Consolidation re-homes a skill's references/ / scripts/ out of the tree
before delete/archive, so the ledger captured only what was left
(files: 1 = SKILL.md) and `hermes curator rollback` restored a hollow
skill — the support files were only recoverable by hand out of the
pre-run .curator_backups tar.

The ledger's delete/archive/purge captures now complete themselves from
the newest curator skills.tar.gz: disk hashes win, the backup fills only
missing paths, tar members escaping the package prefix are rejected,
and every fill target stays under skills/ and HERMES_HOME. The same
fill runs at rollback time, so hollow entries recorded before this fix
still restore the complete package.

Wired at the four capture sites (skill_manage delete, archive_skill,
purge, record_mutation) and verified end-to-end: incident shape
(re-home -> delete -> entry has both files -> rollback restores both),
historical hollow entry repair, no-backup degradation, disk-hash
priority, and tar path-traversal rejection.
2026-08-31 10:08:13 -07:00
Teknium 6fe933e709 fix(cli): launch-context-independent Linux desktop-entry Exec (salvaged from #94874)
Rewrites resolve_exec_command so the generated .desktop Exec no longer
depends on how the installer happened to be launched: fixes the bare
repo-script form whose shebang escapes the venv, and the symlinked-venv
form that .resolve() dereferenced into the base interpreter store.

Salvaged squashed from PR #94874 (24 commits) after the original branch
was found to carry stray __pycache__/.gitignore payload.

Co-authored-by: Gökhan <gkhn.yldrmlr@gmail.com>
2026-08-31 10:07:51 -07:00
Agi-Asi 54ee290bcb fix(dashboard): don't gate Desktop-owned loopback backends on public_url
A non-loopback dashboard.public_url engaged the ticket-only auth gate for
EVERY hermes serve on the machine — including the private loopback
backends the Desktop app spawns for itself (HERMES_DESKTOP=1). Those
backends authenticate with the per-spawn session token, which the gated
WS path refuses outright, so Desktop failed to boot with:

  Local Hermes backend is HTTP-reachable but the WebSocket (/api/ws)
  rejected the session token.

The public_url describes a DIFFERENT deployment: the actual public
dashboard is a separate process on a non-loopback bind whose own startup
keeps its gate. Exempting Desktop-owned loopback backends therefore never
opens the public surface.

Exemption requires ALL of: loopback bind, HERMES_DESKTOP=1 (set by every
Desktop spawn path, local and SSH), and an operator-minted credential
(HERMES_DASHBOARD_SESSION_TOKEN, SSH session token, or owner nonce).
Non-Desktop serves and non-loopback binds keep the exact previous
behaviour — verified by regression tests on both sides of the boundary.

Fixes #96490
2026-08-31 10:07:34 -07:00
chelsealong 6d407ca1a4 fix(desktop): stop model_context_length edits from being dropped or wiped
_denormalize_config_from_web only wrote model_context_length into the
on-disk model dict inside the branch gated on `model` also being present
in the payload. That was harmless when the frontend always sent the full
config, but the prior commit switched Settings autosave to send only the
diff (diffConfig), so editing the Context Window control alone omits
`model` from the payload and the context-length edit is silently thrown
away. The mirror case regressed too: editing `model` alone now omits
model_context_length from the diff, and the old code treated that missing
key the same as an explicit 0, wiping an existing context_length override
that the user never touched.

Track whether model_context_length was actually present in the payload
and only mutate context_length when it was, independent of whether
`model` also changed.
2026-08-31 10:07:17 -07:00
fangliquanflq 915ec169a0 fix(update): verify failed restore cleanup 2026-08-31 10:04:56 -07:00
fangliquanflq 53cf38c3b8 fix(update): fail closed on incomplete restore checks 2026-08-31 10:04:56 -07:00
fangliquanflq 3e4dc5f2c3 fix(update): preserve unknown restore cleanup state 2026-08-31 10:04:56 -07:00
fangliquanflq 8236b51878 fix(update): authenticate import health markers 2026-08-31 10:04:56 -07:00
fangliquanflq 6c608e2f59 fix(update): reject terminated import probes 2026-08-31 10:04:56 -07:00
fangliquanflq 68f9681acb fix(update): capture terminating restored imports 2026-08-31 10:04:56 -07:00
fangliquanflq f6c3942957 fix(update): compare every restored module failure 2026-08-31 10:04:56 -07:00
fangliquanflq 7b958b3575 fix(update): detect restored import-time failures 2026-08-31 10:04:56 -07:00
fangliquanflq 716b10314a fix(update): reject unsafe stash restores 2026-08-31 10:04:56 -07:00
Teknium e49df40682 fix(update): refuse to mutate a venv containing foreign-owned files (#83529)
A venv ever touched by sudo pip / sudo hermes contains root-owned files
(classically site-packages/*.dist-info/INSTALLER). A later normal-user
'hermes update' pulls code fine, then 'uv pip install -e .' dies with
'Permission denied (os error 13)' mid-mutation — venv/bin/hermes already
deleted, CLI bricked.

Add a bounded, pure-stat ownership preflight (_venv_foreign_owned_paths)
that runs after the code pull and immediately before the dependency
install. If foreign-owned paths are found it refuses up front, names the
offending paths + owner uid, prints the exact recovery command
(sudo chown -R $(id -un): <root>), and confirms the venv is untouched.
Windows (no os.geteuid) and root skip entirely. Never raises, capped at
~2000 stat calls, no subprocess use (update tests mock subprocess.run).

Same refuse-before-mutate philosophy as the contended-venv gate (#87331).

Fixes #83529
Diagnosis and documented recovery by @eabase.
2026-08-31 10:04:40 -07:00
Teknium f2f7a3bf15 feat(cron): doctor flags overdue next_run_at as silent non-firing
Widens the salvaged cron doctor with the highest-value fleet check:
an active job whose next_run_at is parked >15min in the past is not
firing (dead ticker, downed gateway, wedged fire-claim). Also registers
doctor in the docs (cron guide + CLI reference) and resolves the salvage
onto current main alongside runs/incidents/notepad.
2026-08-31 10:00:29 -07:00
joe102084 b028fe632e feat: add cron doctor health check 2026-08-31 10:00:29 -07:00
Jay. (neocode24) 9a7732b45f fix(cron): stale ticker yields its tick to a fresh gateway
A long-lived process whose checkout was updated underneath it (hot git
pull, interrupted hermes update) serves mixed sys.modules. When such a
stale process races a fresh gateway for the cron tick lock and wins the
minute, every agent job it dispatches can die on ImportErrors whose real
cause is staleness — and the fresh gateway's ticker skips the same minute
as lock-loser, so the user's scheduled job fires broken or not at all.

tick() now checks, BEFORE acquiring the tick lock:

  skew detected (boot fingerprint != disk revision)
    AND this process does not own the gateway runtime lock
    AND that lock is held (a fresh gateway is alive)
      -> raise CronTickYielded, skipping the tick entirely

Each arm alone keeps the old behavior:
- skew + self-owned lock -> proceed (delivery-path stale-code hint stays
  the surface for gateway-owned dispatches)
- skew + no lock holder -> proceed (desktop-standalone users must not
  lose their only ticker to a silent yield)
- skew None (non-git install, no boot fingerprint, probe failure) ->
  proceed; yielding is a certainty claim, never a guess

The yield RAISES instead of returning 0 so the provider loops record it
via record_ticker_error and mark the heartbeat success=False — a yielded
tick must not look like a healthy one (hermes cron status shows why),
mirroring the EMFILE propagation contract (#87644). Yield logging is
throttled to once per skew episode. Self-healing: when the fresh gateway
dies, its lock releases and the stale ticker's next tick proceeds.

Multiplex loop: a yield for one profile no longer cancels sibling
profiles' ticks in the same cycle; only the yielding profile records an
unsuccessful beat.

gateway/status.py gains owns_gateway_runtime_lock() —
is_gateway_runtime_lock_active() is True for the lock's own owner too, so
a caller deciding whether to yield to a FRESH gateway needs the
in-process handle as the discriminator.
2026-08-31 09:59:07 -07:00
Teknium fc2421cfa7 fix(gateway): hold inbound gate until turn machinery is warm on fresh boot (#99373)
On a fresh boot with no resume_pending sessions, _finish_startup_restore
opened the inbound gate almost immediately while the agent-side turn
machinery (run_agent import graph, tool schemas + check_fn probes,
context-file tier) was still cold. A message arriving in that window was
served with a skeleton system prompt (~1.7K tokens vs ~14.6K healthy):
no AGENTS.md/context tier, no tool schemas, memory provider initializing
mid-turn.

Fix: start a background turn-machinery warm-up when the startup gate
closes (overlapping the network-bound platform connects) and have
_finish_startup_restore await it — BOUNDED by
agent.gateway_startup_warmup_timeout (default 20s, 0 disables) — before
draining the queue and opening the gate. On timeout the gate opens
anyway and the warm-up finishes in the background, so a wedged init can
never make the gateway permanently unavailable.

Reported by @yhfmstr in #99373.

Fixes #99373
2026-08-31 09:58:24 -07:00
Nio Thomas ba7743b076 fix(state-db): report corruption instead of "session not found", detect it early
Re-applied onto 3aee29089 after `hermes update` reset main to origin/main.

1. web_routers/sessions.py: _resolve_session_id() classifies malformed-DB
   errors via the existing is_malformed_db_error() and raises 503 at all five
   call sites. delete_session_endpoint was the worst — an unresolvable id
   counted as idempotent success, so DELETE reported it had removed a session
   that was still on disk.
2. gateway/lifecycle_ledger.py: check_state_db_integrity() runs PRAGMA
   quick_check(1) on the unclean-exit path only (~2s on 500MB) and records the
   verdict into gateway-exit-diag.log. The 2026-08-31 corruption sat undetected
   for 3.5 days because nothing ever looked.
3. hermes_cli/gateway.py: `gateway run --replace` gave the outgoing gateway 5s
   before SIGKILL; SessionDB.close() runs a PASSIVE WAL checkpoint that does
   not finish in 5s on a WAL 4x past the autocheckpoint threshold, and a kill
   mid-checkpoint tears b-tree pages. Grace raised to 30s via a testable
   _await_gateway_exit() that also re-checks after the final sleep (a PID
   exiting in the last interval must not be SIGKILLed — PID-reuse hazard).

NOT added: wal_checkpoint(TRUNCATE) at shutdown — removed upstream in #45383
because a TRUNCATE reset races the live writer and tears b-tree pages.

Adversarial review: Codex gpt-5.6-sol, 9.0/10 across three groups, no must-fix.
2026-08-31 09:56:54 -07:00
Teknium cdd3b84c13 fix(state): reject special files in zeroed probe; real schema-bytes decode fixture
Follow-ups on the salvage: regular-file guard before the zeroed byte-probe (a FIFO at the state.db path would block startup forever — #98017 review P2), plus an on-main-reproducing UnicodeDecodeError fixture for #98924 (raw bytes in sqlite_master, not messages.content, are what reach pysqlite error-message decode).
2026-08-31 09:56:43 -07:00
Alvin T. Veroy e17fd0a708 fix(state): decode errors now reach the heal path and fail loud in TUI (residual #98924 surfaces)
Companion to #98935, which fixes _fts_table_probe itself. This covers the
surfaces that PR does not touch:

- web_server._open_session_db_at_path: the one-writable-open heal only
  caught sqlite3.DatabaseError; a raw UnicodeDecodeError (pysqlite failing
  to decode SQLite's own error message over corrupt file bytes) bypassed
  it, so the heal documented for malformed schema never fired (#98924
  Failure 1). Both catches widened; decode errors dispatch to the heal.
- SessionSchemaMixin._recover_stale_fts_locked: drop-and-recreate skipped
  vtables whose probe raised UnicodeDecodeError, the same too-narrow
  catch the issue identified in the probe.
- TUI gateway: _ensure_session_db_row returned silently when the store
  could not open, so prompt.submit streamed the turn while persisting
  nothing (#98924 Failure 2). It now returns False and prompt.submit
  fails the RPC with code 5072 so desktop maps it to a toast, mirroring
  the disk-full/5070 convention. session.create stays silent per its
  pinned degraded-mode contract.
2026-08-31 09:56:43 -07:00
loulanyue 69245e65bd fix(state): serialize startup across zero-byte check, quarantine, connect, and schema commit (#97568)
- Guard against concurrent-opener race where newly created 0-byte state.db was falsely quarantined before first schema write
- Wrap startup in quarantine_cross_process_lock when database is uninitialized or zeroed
- Guard is_zeroed_sqlite_file and is_zeroed_state_db against active live connections in current process
- Add concurrent-opener and live-connection regression tests
2026-08-31 09:56:43 -07:00
kshitijk4poor 936b970e28 fix(browser): lightpanda review follow-ups for #99312
- lightpanda_engine_status: check use_real_profile before the cloud
  provider, matching browser_exec's actual resolution order (real-profile
  resolution runs before backend resolution), so /browser status and
  hermes doctor name the right shadowing setting when both are set.
- launch_lightpanda: drop the unreachable Windows popen_kwargs branch
  (find_lightpanda_binary returns None on nt, launch errors out earlier).
- doctor: drop the over-defensive try/except around the cached
  _using_lightpanda_engine() config read.
- Docstring: 'no-I/O gates' -> 'no network I/O (config reads only)'.
- New test pinning real-profile-over-cloud-provider reason precedence.
2026-08-31 21:59:11 +05:30
Adrià Arrufat 8bdf8836e6 docs(browser): document Lightpanda in Browser Use mode and the engine precedence rules 2026-08-31 21:59:11 +05:30
Adrià Arrufat e3a85ae5a0 feat(browser): honor browser.engine=lightpanda in Browser Use mode
Browser Use mode never read browser.engine: _resolve_backend_cdp() went
BU_CDP_* env -> CDP override -> cloud provider -> local Chrome, so
`engine: lightpanda` was a silent no-op on the default backend, and on
the built-in path it was skipped whenever a cloud provider, Camofox or a
CDP override was active without anyone saying so.

- browser_use_cli: when the engine is lightpanda and nothing with higher
  precedence claimed the session, get a session from _get_session_info()
  and export its endpoint as BU_CDP_URL; the browser is private to the
  session key, so the own-tab preamble is skipped. The browser_exec
  description gains a Lightpanda header (text-first, new_tab once then
  goto_url — lightpanda-io/browser#1962).
- browser_tool: _create_local_session() spawns `lightpanda serve
  --host 127.0.0.1 --port <free>` per session key (new
  tools/browser_lightpanda.py), reusing the session cache, inactivity
  reaper and atexit cleanup; a dead process is respawned on the next call;
  orphans from a crashed Hermes are reaped through per-process records in
  $HERMES_HOME/cache/browser-use/lightpanda/. New lightpanda_engine_status()
  reports whether the engine is in effect or what shadows it.
- tools_config: "Lightpanda" row in the Browser Automation picker
  (cloud_provider: local + engine: lightpanda; "Local Browser" resets the
  engine to auto) with a binary-check post-setup.
- /browser status and hermes doctor print the engine state and, when it
  is shadowed, the reason.
2026-08-31 21:59:11 +05:30
Teknium ca9952cbb3 fix(plugins): doctor temp home can no longer be stranded by a failed staging copy
Follow-up to the salvaged manifest guard (#90859): the doctor's temporary
HERMES_HOME now enters the ExitStack before the staging copytree, so
ENOSPC / KeyboardInterrupt / any exception during staging deterministically
removes the hermes-plugin-doctor-* directory instead of relying on the
TemporaryDirectory GC finalizer (which cannot run while the traceback pins
the frame). Regression test proven via sabotage run against the old code.
2026-08-31 08:41:15 -07:00
DevKJ a00ae1d519 fix(cli): stop plugins doctor from copying a non-plugin directory
`hermes plugins doctor` defaults its target to `.`, and
resolve_plugin_path accepted any directory that existed. Doctor then
copytree'd the resolved path into a temporary HERMES_HOME *before* the
runtime got to reject it, so running the command from a directory that
is not a plugin copied that whole tree.

Run from $HOME on macOS this copies the home directory, and because
`Library/CloudStorage` is not excluded it also materializes every
cloud-only Google Drive/iCloud placeholder. Observed locally: 461 GB
written to /private/var/folders and still growing when the process was
killed, on a machine with 49 GB free.

Resolve now requires a manifest before returning a path, mirroring
PluginManager._scan_directory: `plugin.yaml`/`plugin.yml`/`plugin.json`
in the directory itself, or in one immediate subdirectory for the
category layout. Plugin-id candidates are only tried for values that can
be an id, since joining `.` onto a plugins root resolves to the root and
would hand Doctor every installed plugin at once.

Non-plugin targets now fail with a clear message and no copy.
2026-08-31 08:41:15 -07:00
Kshitij Kapoor 9488950a9d fix(gateway): build reaper exclusion from raw registration records
/simplify-code efficiency reviewer (verified): cleanup_stale=False does
NOT deliver the exclusion the salvaged fix intended — get_running_pid
returns None whenever a record fails liveness VALIDATION (start-time
mismatch, argv drift, lock hiccup) regardless of the flag, which only
controls unlinking. In exactly the at-risk scenario the recorded PID
still never joined the exclusion set.

Exclusion evidence now comes from the RAW pidfile + lock records (no
validation, no unlink side effects); the validated non-destructive
probe is kept for the runtime-status fallback PID. For a KILL exclusion
list this is strictly safer: a stale recorded PID at worst spares one
process for one sweep, while a validation false-negative would
TerminateProcess a live gateway. Regression test reworked to drive the
real function semantics (raw record present, validation rejects);
mutation-checked: removing the raw-record read fails the test.
2026-08-31 21:09:43 +05:30
Kshitij Kapoor c26762d60e fix(desktop): keep the orphan sweep unconditional; guard the probe fix with a mutation-checked test
Follow-up on the salvaged #87158: the reaper-side fix (probe with
cleanup_stale=False so the sweep never deletes its own exclusion
evidence) fully protects a healthy standalone gateway, so the
web_server.py skip-sweep gate is dropped — a stale-but-present
registration must not veto the #77276 orphan reap that motivated the
sweep in the first place.

New regression test drives the exact failure: a registration that
fails liveness validation only surfaces its PID when probed
non-destructively; with the destructive default the standalone gateway
would be hard-killed (TerminateProcess, no drain). Mutation-checked:
reverting the probe to get_running_pid() fails the test.
2026-08-31 21:09:43 +05:30
butbutbutbutbutbut 1ec4b569f2 fix(desktop): don't reap the healthy standalone gateway on Windows desktop startup 2026-08-31 21:09:43 +05:30
loulanyue a071fc80da fix(state): quarantine 0-byte truncated state.db and record store provenance (#97568) 2026-08-31 20:44:26 +05:30
686f6c61 e703717513 fix(cron): treat a live multiplexer as gateway-alive for satellite profiles
A named profile has no local gateway.pid, so cron warned that jobs
would not fire and recommended hermes gateway install — which the
start guard then refuses with exit 78. Share the multiplexer-serving
probe with the start guard and count it as liveness.
2026-08-31 07:28:39 -07:00
Wesley Simplicio 8edaa25746 fix(cron): isolate desktop profile persistence 2026-08-31 07:28:22 -07:00
Teknium 3a351a9665 feat(worktree): pushed open-PR lanes reclaim their disk; cron tick prunes worktrees
Two growth leaks closed:

1. Pushed-branch tier (the dominant survivor class — 24 of 33 preserved
   trees, ~18GB on the reporting box): managed installs fetch with a
   single-branch refspec, so pushed PR branches never get refs/remotes/*
   entries and read as 'unpushed' forever. When a clean tree's branch head
   EXACTLY matches origin (one lazy git ls-remote per sweep), the checkout
   is redundant: reap the TREE, keep the BRANCH ref (shielded from the
   orphaned-branch pass). Anything diverged/unverifiable stays preserved.
   Applied to both the startup pruner and hermes worktree prune/list.

2. Cron-tick maintenance: the pruner only ran on hermes -w launches, so
   gateway-driven boxes accumulated trees for days. The scheduler tick now
   dispatches the same conservative pruner on a daemon thread, throttled
   to once per 6h, against the install checkout + job-workdir repos that
   have a .worktrees/ dir.
2026-08-31 07:27:50 -07:00
Teknium 9f8557a09f fix(cli): update _get_service_pids docs for profile-scoped systemd filtering 2026-08-31 06:02:32 -07:00
kokhlo 82d7a13003 fix(cron): profile isolation — systemd filter + heartbeat guard
Repairs #98790 where
✓ Gateway is running — cron jobs will fire automatically
  PID: 4165
  Ticker heartbeat: 39s ago

  4 active job(s)
  Next run: 2026-08-30T22:50:18.762041+03:00 in profile B incorrectly reports
that jobs will fire based on profile A's gateway process.

Root causes:
1.  executed ,
   which enumerated the entire systemd fleet regardless of ,
   violating the docstring "only PIDs belonging to the current profile".
2.  checked  when
   (no heartbeat file) should trigger a warning — instead, it fell through
   to the "✓ Gateway is running" green branch.

Changes:
- hermes_cli/gateway.py::_get_service_pids: pattern = get_service_name()
  when all_profiles=False, filtering to the current profile's systemd unit.
- hermes_cli/cron.py::cron_status: guard hb_age is None first with an
  explicit yellow warning: "ticker has not reported a heartbeat".

Regression test suite guards both systemd scoping (default + all_profiles)
and heartbeat branching (None vs fresh vs stale).
2026-08-31 06:02:32 -07:00
teknium1 8fd144c502 fix(desktop): model assignment carries the credential pointer, not a resolved key (#88990, salvage #90484)
Upgrades yesterday's #99310 skip-guard to full pointer-carry from
PR #90484: model assignment and custom-endpoint activation now write
key_env or the raw ${VAR} template into model config instead of
dropping the credential reference entirely, so the model entry keeps
resolving at runtime with zero plaintext in config.yaml. Applied
surgically onto current main (the PR branch predates newer
web_server.py changes); key_env carry made independent of the
expanded api_key guard, tests updated to pin pointer-carry.
2026-08-31 04:40:24 -07:00
Teknium a90be562f4 fix(web): stop mirroring env-backed provider keys into model.api_key (#88990)
POST /api/model/set copied the load_config()-resolved plaintext of a
${VAR}/key_env provider entry into model.api_key, writing the secret
into config.yaml and recreating it on every re-apply. The mirror now
checks the RAW on-disk entry and skips env-referencing entries;
literal keys keep the existing behavior.
2026-08-31 03:37:43 -07:00
Frowtek 1152d4d3ce fix(cli): recognize whitespace around '=' in .env save/remove
_env_line_defines_key() decides which .env lines the writers may rewrite or
drop. It matched on the `KEY=` prefix, but load_env() splits on the first
`=` and strips the name:

    key, _, value = line.partition('=')
    env_vars[key.strip()] = _parse_env_value(value)

so `OPENAI_API_KEY = sk-...` is a live assignment — the key resolves, the
provider works, and every UI shows it as set. The writers did not see it.

This is the same resurrection hole #40041 fixed for `export KEY=`, still
open for the whitespace form:

- DELETE /api/env 404s ("not found in .env") while the credential stays
  active — a key the user revoked through the UI is never actually revoked
- PUT /api/env appends a SECOND line instead of replacing; a later delete
  removes the appended line and the original value silently comes back

Rotate-then-delete on a spaced line therefore restores exactly the key the
user rotated away from.

Match load_env()'s parse instead of prefix-matching, so the writers accept
precisely what the reader accepts: skip blank/comment/no-'=' lines, strip an
`export ` prefix, then compare the stripped name. Commented-out lines stay
untouched and `KEY_EXTRA=`/`MY_KEY=` still do not match `KEY`.

Verified against the real dashboard endpoints on a temp HERMES_HOME: the
spaced line is now removed, rotation replaces it in place with no duplicate,
and a parity check asserts the writer matches a line iff load_env() does.
2026-08-31 03:37:43 -07:00