Follow-up on top of @rahlquist's terminal.temp_dir knob (#97182): the
default itself now avoids RAM-backed tmpfs. Resolution order on the
local backend: terminal.temp_dir > TMPDIR/TMP/TEMP > HERMES_HOME/cache/
terminal (managed, pruned) > /tmp fallback. Pruning: hourly via gateway
housekeeping + once-per-process best-effort sweep; hermes_bg_* triplets
are aged as a group so a live server's fresh .log protects its .pid.
Two sibling readers of raw config.mcp_servers duplicated their own (weaker)
shape guard: mcp-health.ts guarded the map but still passed null entries to
isUrlServer (crash on .url read), and the command palette re-implemented the
map check inline. Both now go through getServers(), the single choke point
that drops malformed entries, so a null entry can't crash the sweep and the
palette lists exactly the servers the MCP tab shows.
Also records the contributor email mapping for the salvaged commits.
Follow-up to the cherry-picked #94338.
Treat the ChatGPT Codex invalid image-data 400 as an image rejection so Hermes strips image parts and retries text-only instead of aborting the session. Add coverage for the exact error wording.
Bot Chats created before the follow_profile_config marker existed carry no
contract in model_config, so they would stay pinned to a stale stored
provider until deleted — the exact shape of the live reports (#89497,
#94818). Mirror the plugin's own identity rule (the profile's session
titled exactly 'Bot Chat') as a legacy fallback in
_stored_session_runtime_overrides, matching the room-plumbing legacy
'Group:' title fallback.
Follow-up to the salvaged #90343 (@curator8888) and #96111 (@lorzl).
- The stale tick-socket sweep's os.kill(pid, 0) liveness probe sits inside
an explicit os.name == 'posix' gate (AF_UNIX nodes never exist on
Windows) — suppress with the standard inline marker.
- contributors/emails/: rodrigo.smscom@gmail.com -> rodrigogs (author of
the salvaged #92315 commits), unblocking check-attribution.
A live sibling serve sharing state.db is no longer treated as a dead process by the startup orphan sweep.
Covers the sweep half of #94895. The launchd Errno 48 KeepAlive loop is not addressed here.
Credit: @Finn763
_sqlite_connect opened a connection via connect_tracked and then ran the
busy_timeout PRAGMA; if that raised, the half-open connection was abandoned
— leaking its fd AND leaving a stale entry in the sqlite_safe_read
live-connection registry (which only clears on close), permanently blocking
byte-level probes of the kanban database. Close before re-raising.
Salvaged from PR #96290 (kanban slice) with regression test.
Follow-ups to the salvaged #95947 cron commit:
- Wrap the lock probe in its own try/except: a crashing probe is
'unknown', not 'dead' — the pid scan still decides instead of the
whole tri-state collapsing to None.
- Regression tests (shape adapted from #94155 by @liuhao1024): lock
held + empty pid scan -> alive (the reported false alarm); lock
inactive -> pid-scan fallback both ways; crashing lock probe still
falls back.
- patch_liveness now pins the lock probe inactive by default so the
pre-existing pid-scan tests stay deterministic on machines where a
real gateway holds the real lock.
- contributors mapping for magnus.lundstedt@infidyne.com.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
The script-timeout path used a site-local process-group kill, which
cannot reach a grandchild that created its OWN session (start_new_session
background jobs, watchdogs). Such descendants kept running after the job
reported failure (#71148, #59549). Migrate the timeout handler to the
unified deadline layer's kill_process_tree (#85147, d6a5cb9725): psutil
snapshots the descendant set before signalling, so own-session
grandchildren are reached too. Fallback to the site-local group kill if
the import ever fails, so the path cannot re-wedge.
The explicit script-timeout message stays the classification anchor
(#85536's contract), keeping cron timeouts distinct from provider
timeouts.
Salvage additions on review (#85125 Phase 4a):
- migrate the sibling kill site too — the cancel_event/"ownership was
lost" path orphaned setsid grandchildren the same way (whole-bug-class
rule); pinned by test_cancel_path_also_tree_kills
- proc.poll() early-return in _terminate_cron_script_tree so a script
that exits right at the deadline doesn't log a spurious "no signal"
warning (mirrors _terminate_cron_script_process); pinned by
test_already_exited_proc_is_left_alone
- acceptance test's script timeout 1s -> 2s: interpreter startup under
CI load could eat the whole 1s window before the spawner wrote its
pid file
- note: kill_process_tree hard-kills (SIGKILL) immediately, whereas the
old path gave a 1s SIGTERM grace window; intended for a deadline-
expiry hard stop (both docstrings say "hard stop")
Based on #86791 by @ayushnangia; cherry-picked to preserve authorship.
Co-authored-by: dante32683 <dante32683@users.noreply.github.com>
Co-authored-by: supotato-ipj <supotato-ipj@users.noreply.github.com>
On Windows installs where the gateway runs as an SCM service (WinSW,
NSSM, sc.exe create), the existing pause machinery kills the gateway
process directly — and the service wrapper's failure ladder resurrects
it within seconds, re-taking the venv file locks mid-update. The update
then dies partway through dependency sync with access-denied errors.
This extends _pause_windows_gateways_for_update() to detect when a
gateway's process tree is owned by a running SCM service, and to stop
the SERVICE through sc.exe instead of killing the child:
- gateway/status.py: expose service-ownership discovery for gateway
runtimes (find_windows_gateway_services maps validated gateway PIDs
through process ancestry to running SCM service PIDs, with
create-time identity checks against PID reuse).
- hermes_cli/update_cmd.py: stop verified services via sc.exe before
venv mutation and restart them afterward. Stops wait for a stable
SCM 'stopped' state AND for the original descendant processes to
exit (service 'Stopped' is not proof the child released its
handles). Failure to prove ownership, stop a service, or restart it
fails closed; rollback restores attempted services, and rollback
failures are surfaced rather than swallowed.
- Fail-closed throughout: unreadable identities, ambiguous ancestry,
or a service that will not reach a stable state abort the update
before any file mutation.
Complements #37039 (gateway-only concurrent instances no longer abort):
that fix lets the update proceed past the gate; this one makes the
pause actually stick when the gateway is service-supervised.
Note: tests/gateway/test_status.py::TestReadProcessCmdlinePsFallback::
test_ps_fallback_when_proc_unavailable fails on Windows on current main
before this change as well (POSIX ps fallback asserted on a platform
without it); all other touched suites pass (155 passed, 5 skipped).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The post-update fleet version check slept 2s and probed once. On Windows the
resume path relaunches the gateway detached, and it needs ~10s to boot (the
Telegram polling reconnect) before it stamps gateway_state.json or answers the
control socket. That race reported "no rows" for a healthy resume, exited 1,
and triggered a full retry that re-killed the gateway the first attempt had
just started — leaving it down and surfacing "Update failed (exit 1)".
Poll a bounded window (up to 30s) for the resumed gateway to publish its
identity, and only treat a persistently empty snapshot as verification
failure. The fail-closed contract from #93406 is preserved: a gateway that
genuinely never comes back still exits 1.
mount_spa's WEB_DIST.exists() check ran ONCE at mount time: a long-lived
'hermes dashboard --skip-build' that survived a git pull (or launched
before the first build) installed a permanent no_frontend catch-all and
answered 404 'Frontend not built' on every route forever — even after
npm run build completed. Remote Desktop clients saw ERR_EMPTY_RESPONSE.
The missing-dist branch is now reserved for the headless-serve contract
only. The SPA routes mount unconditionally and already cope with a
missing dist per-request (_serve_index returns the same 404 JSON when
index.html is unreadable; the /assets mount gains check_dir=False so
StaticFiles 404s instead of raising at mount). The dashboard recovers
the moment a build lands on disk — no restart needed.
Direction from #82666 by @codexbt (his PR's rebase dropped the product
hunk, leaving only the test; the test is cherry-picked as-is and this
commit restores the behavior it pins, adapted to the current mount_spa
shape: headless guard preserved, per-request recovery instead of a
per-request exists() check).