Commit Graph

695 Commits

Author SHA1 Message Date
Teknium c5b44e0756 chore: map contributor emails for TiberiuD and fedebyes 2026-08-28 07:51:23 -07:00
Teknium 3f3ae6850a chore: map brianbaldock contributor email 2026-08-28 07:51:16 -07:00
Teknium 95cf7dc9e8 feat: session temp root moves off tmpfs /tmp to ~/.hermes/cache/terminal by default; auto-pruned after 72h
Follow-up on top of @rahlquist's terminal.temp_dir knob (#97182): the
default itself now avoids RAM-backed tmpfs. Resolution order on the
local backend: terminal.temp_dir > TMPDIR/TMP/TEMP > HERMES_HOME/cache/
terminal (managed, pruned) > /tmp fallback. Pruning: hourly via gateway
housekeeping + once-per-process best-effort sweep; hermes_bg_* triplets
are aged as a group so a live server's fresh .log protects its .pid.
2026-08-28 07:50:33 -07:00
Teknium 9f2ab334c8 fix(desktop): route the MCP health sweep and command palette through getServers
Two sibling readers of raw config.mcp_servers duplicated their own (weaker)
shape guard: mcp-health.ts guarded the map but still passed null entries to
isUrlServer (crash on .url read), and the command palette re-implemented the
map check inline. Both now go through getServers(), the single choke point
that drops malformed entries, so a null entry can't crash the sweep and the
palette lists exactly the servers the MCP tab shows.

Also records the contributor email mapping for the salvaged commits.

Follow-up to the cherry-picked #94338.
2026-08-28 06:33:40 -07:00
Teknium 911a41ec97 chore: map contributor email for hakanbaysal 2026-08-28 05:17:26 -07:00
fangliquanflq bb06a8d414 fix(dashboard): harden Node engine alignment checks 2026-08-28 05:12:33 -07:00
Teknium 64caddc06a chore: contributor mapping for HaiyiMei 2026-08-28 04:58:21 -07:00
Teknium b0f40ade4d chore: contributor mapping for ahrazzle 2026-08-28 04:57:49 -07:00
Teknium 7aebf300a7 chore: contributor mapping for fkdls112 2026-08-28 03:46:24 -07:00
Al Cooke 0241619068 fix: retry text-only on Codex invalid image data errors
Treat the ChatGPT Codex invalid image-data 400 as an image rejection so Hermes strips image parts and retries text-only instead of aborting the session. Add coverage for the exact error wording.
2026-08-28 03:46:24 -07:00
Teknium af53d02920 fix(bot-mode): backfill follow-profile contract for legacy canonical Bot Chats
Bot Chats created before the follow_profile_config marker existed carry no
contract in model_config, so they would stay pinned to a stale stored
provider until deleted — the exact shape of the live reports (#89497,
#94818). Mirror the plugin's own identity rule (the profile's session
titled exactly 'Bot Chat') as a legacy fallback in
_stored_session_runtime_overrides, matching the room-plumbing legacy
'Group:' title fallback.

Follow-up to the salvaged #90343 (@curator8888) and #96111 (@lorzl).
2026-08-28 02:59:17 -07:00
kshitijk4poor 06141577cd chore: AUTHOR_MAP add mrkillbob (mikedemott@Mikes-Mac-mini.local)
Contributor email for PR #93634 salvage. GitHub user mrkillbob (ID 25466867).
2026-08-28 12:38:49 +05:30
kshitijk4poor f4ae8958b5 chore: map zengzheqing noreply email for attribution (PR #96647) 2026-08-28 09:05:19 +05:30
kshitijk4poor 28ee6ac043 chore: map contributor emails for attribution (#95433, #94996) 2026-08-28 02:29:27 +05:30
Teknium 9522c4e8e2 chore: map contributor email for salvage of #95956 2026-08-27 13:48:45 -07:00
Teknium d91e4376f1 Merge pull request #65108 from NousResearch/hermes/hermes-793f4fd9
feat(skills): rewrite AgentMail optional skill CLI-first (salvages #60811)
2026-08-27 12:45:01 -07:00
kshitijk4poor f6f707b783 chore: suppress posix-gated os.kill probe in footgun scan; map rodrigogs in AUTHOR_MAP
- The stale tick-socket sweep's os.kill(pid, 0) liveness probe sits inside
  an explicit os.name == 'posix' gate (AF_UNIX nodes never exist on
  Windows) — suppress with the standard inline marker.
- contributors/emails/: rodrigo.smscom@gmail.com -> rodrigogs (author of
  the salvaged #92315 commits), unblocking check-attribution.
2026-08-27 22:06:17 +05:30
Finn763 cae58be1f5 fix(state.db): cross-backend heartbeat gates orphan sweep
A live sibling serve sharing state.db is no longer treated as a dead process by the startup orphan sweep.

Covers the sweep half of #94895. The launchd Errno 48 KeepAlive loop is not addressed here.

Credit: @Finn763
2026-08-27 11:31:59 -05:00
kshitij 59805b1dd2 chore: map contributor email for deadczarvc 2026-08-27 21:36:35 +05:30
kshitijk4poor 5dc2049bba chore: map joby@ellingtonlife.com -> gijoby (PR #91354 salvage) 2026-08-27 20:06:04 +05:30
Teknium fef1790165 chore: map contributor email for @crazyhulk 2026-08-27 07:33:36 -07:00
Teknium 2d0d2e4105 chore(contributors): map Haakam Aujla email for #60811 salvage 2026-08-27 07:11:52 -07:00
Heath Harris a71be9852e fix(kanban): close half-open tracked connection when busy_timeout PRAGMA fails
_sqlite_connect opened a connection via connect_tracked and then ran the
busy_timeout PRAGMA; if that raised, the half-open connection was abandoned
— leaking its fd AND leaving a stale entry in the sqlite_safe_read
live-connection registry (which only clears on close), permanently blocking
byte-level probes of the kanban database. Close before re-raising.

Salvaged from PR #96290 (kanban slice) with regression test.
2026-08-27 19:07:15 +05:30
kshitijk4poor 3516438832 chore(contributors): map me@kitsonkelly.com -> kitsonk 2026-08-27 17:11:56 +05:30
Teknium a588685fb9 chore(contributors): map aydinhrrs@gmail.com -> RibatTRW 2026-08-27 03:57:25 -07:00
kshitijk4poor f2d043eb08 test(cron): pin lock-first liveness + harden lock-probe failure
Follow-ups to the salvaged #95947 cron commit:

- Wrap the lock probe in its own try/except: a crashing probe is
  'unknown', not 'dead' — the pid scan still decides instead of the
  whole tri-state collapsing to None.
- Regression tests (shape adapted from #94155 by @liuhao1024): lock
  held + empty pid scan -> alive (the reported false alarm); lock
  inactive -> pid-scan fallback both ways; crashing lock probe still
  falls back.
- patch_liveness now pins the lock probe inactive by default so the
  pre-existing pid-scan tests stay deterministic on machines where a
  real gateway holds the real lock.
- contributors mapping for magnus.lundstedt@infidyne.com.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-08-27 16:10:57 +05:30
Teknium 74ec63fe18 chore: contributor email mapping for ekzhang 2026-08-27 03:24:20 -07:00
Ayush Nangia 2f0f01192d fix(cron): tree-kill script timeout descendants via agent.deadline.kill_process_tree
The script-timeout path used a site-local process-group kill, which
cannot reach a grandchild that created its OWN session (start_new_session
background jobs, watchdogs). Such descendants kept running after the job
reported failure (#71148, #59549). Migrate the timeout handler to the
unified deadline layer's kill_process_tree (#85147, d6a5cb9725): psutil
snapshots the descendant set before signalling, so own-session
grandchildren are reached too. Fallback to the site-local group kill if
the import ever fails, so the path cannot re-wedge.

The explicit script-timeout message stays the classification anchor
(#85536's contract), keeping cron timeouts distinct from provider
timeouts.

Salvage additions on review (#85125 Phase 4a):
- migrate the sibling kill site too — the cancel_event/"ownership was
  lost" path orphaned setsid grandchildren the same way (whole-bug-class
  rule); pinned by test_cancel_path_also_tree_kills
- proc.poll() early-return in _terminate_cron_script_tree so a script
  that exits right at the deadline doesn't log a spurious "no signal"
  warning (mirrors _terminate_cron_script_process); pinned by
  test_already_exited_proc_is_left_alone
- acceptance test's script timeout 1s -> 2s: interpreter startup under
  CI load could eat the whole 1s window before the spawner wrote its
  pid file
- note: kill_process_tree hard-kills (SIGKILL) immediately, whereas the
  old path gave a 1s SIGTERM grace window; intended for a deadline-
  expiry hard stop (both docstrings say "hard stop")

Based on #86791 by @ayushnangia; cherry-picked to preserve authorship.

Co-authored-by: dante32683 <dante32683@users.noreply.github.com>
Co-authored-by: supotato-ipj <supotato-ipj@users.noreply.github.com>
2026-08-27 14:28:06 +05:30
kshitij f45477b4a8 Merge pull request #86412 from kshitijk4poor/salvage/83225-overflow-clamp
fix(approval): oversized approvals.timeout crashes parallel tool batches — clamp at config read (salvage #83225/#83298, #85125 2b)
2026-08-27 14:26:38 +05:30
kshitijk4poor 5d4ad23b0d chore: map paulapsp157@gmail.com -> Aoshi-Dev (PR #96003 salvage) 2026-08-27 14:02:06 +05:30
Victor Kyriazakos db00793db7 chore(contributors): map potatosaladx@gmail.com to potatosalad (attribution for salvaged #79436/#81128 commits) 2026-08-27 15:33:32 +10:00
Teknium 26e48dd8c3 chore: map contributor email for kvnloo (#95173) 2026-08-26 21:38:20 -07:00
Teknium 89f32fe4a5 chore: map fred0m noreply email in contributors registry (#91360 salvage) 2026-08-26 17:42:50 -07:00
Casey 790e1eb6bd fix(update): pause SCM-supervised Windows gateway services before venv mutation
On Windows installs where the gateway runs as an SCM service (WinSW,
NSSM, sc.exe create), the existing pause machinery kills the gateway
process directly — and the service wrapper's failure ladder resurrects
it within seconds, re-taking the venv file locks mid-update. The update
then dies partway through dependency sync with access-denied errors.

This extends _pause_windows_gateways_for_update() to detect when a
gateway's process tree is owned by a running SCM service, and to stop
the SERVICE through sc.exe instead of killing the child:

- gateway/status.py: expose service-ownership discovery for gateway
  runtimes (find_windows_gateway_services maps validated gateway PIDs
  through process ancestry to running SCM service PIDs, with
  create-time identity checks against PID reuse).
- hermes_cli/update_cmd.py: stop verified services via sc.exe before
  venv mutation and restart them afterward. Stops wait for a stable
  SCM 'stopped' state AND for the original descendant processes to
  exit (service 'Stopped' is not proof the child released its
  handles). Failure to prove ownership, stop a service, or restart it
  fails closed; rollback restores attempted services, and rollback
  failures are surfaced rather than swallowed.
- Fail-closed throughout: unreadable identities, ambiguous ancestry,
  or a service that will not reach a stable state abort the update
  before any file mutation.

Complements #37039 (gateway-only concurrent instances no longer abort):
that fix lets the update proceed past the gate; this one makes the
pause actually stick when the gateway is service-supervised.

Note: tests/gateway/test_status.py::TestReadProcessCmdlinePsFallback::
test_ps_fallback_when_proc_unavailable fails on Windows on current main
before this change as well (POSIX ps fallback asserted on a platform
without it); all other touched suites pass (155 passed, 5 skipped).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 16:45:31 -07:00
Finn763 de34c746cc chore: map contributor emails 2026-08-26 15:51:22 -07:00
AlexGabbia b3e477f304 fix(update): wait for resumed Windows gateway before failing fleet check
The post-update fleet version check slept 2s and probed once. On Windows the
resume path relaunches the gateway detached, and it needs ~10s to boot (the
Telegram polling reconnect) before it stamps gateway_state.json or answers the
control socket. That race reported "no rows" for a healthy resume, exited 1,
and triggered a full retry that re-killed the gateway the first attempt had
just started — leaving it down and surfacing "Update failed (exit 1)".

Poll a bounded window (up to 30s) for the resumed gateway to publish its
identity, and only treat a persistently empty snapshot as verification
failure. The fail-closed contract from #93406 is preserved: a gateway that
genuinely never comes back still exits 1.
2026-08-26 15:05:32 -07:00
Teknium 3a4c309cfb chore: map contributor email for salvage of #88470 2026-08-26 10:10:11 -07:00
Teknium 4e8e5f84bc chore: contributor mappings for #95633/#95652 salvage 2026-08-26 09:51:55 -07:00
etzelvon dbbed456c1 style(desktop): satisfy curly + padding-line lint for guarded switch
- braces for single-statement if(confirmNotificationId) guards
- blank line before return in staleness branch (padding-line-between-statements)
2026-08-26 08:39:45 -07:00
Teknium f4e8eb1568 chore: map A-Snegin attribution 2026-08-26 07:28:10 -07:00
Teknium 1145fcaed6 chore: map contributor email for deepeet-git 2026-08-26 06:26:02 -07:00
Teknium 279726cc2f chore: map contributor email for wiseconnex 2026-08-26 04:51:41 -07:00
Teknium 6bbae974d5 chore: map justinjohnson25600 and BrunoBza contributor emails 2026-08-26 04:49:22 -07:00
Teknium 20d33e385a fix(web_server): a dashboard started without a build recovers the moment one appears (#82614)
mount_spa's WEB_DIST.exists() check ran ONCE at mount time: a long-lived
'hermes dashboard --skip-build' that survived a git pull (or launched
before the first build) installed a permanent no_frontend catch-all and
answered 404 'Frontend not built' on every route forever — even after
npm run build completed. Remote Desktop clients saw ERR_EMPTY_RESPONSE.

The missing-dist branch is now reserved for the headless-serve contract
only. The SPA routes mount unconditionally and already cope with a
missing dist per-request (_serve_index returns the same 404 JSON when
index.html is unreadable; the /assets mount gains check_dir=False so
StaticFiles 404s instead of raising at mount). The dashboard recovers
the moment a build lands on disk — no restart needed.

Direction from #82666 by @codexbt (his PR's rebase dropped the product
hunk, leaving only the test; the test is cherry-picked as-is and this
commit restores the behavior it pins, adapted to the current mount_spa
shape: headless guard preserved, per-request recovery instead of a
per-request exists() check).
2026-08-26 04:48:58 -07:00
Teknium 7a10d91b29 chore: map contributor email for notkisk 2026-08-26 04:14:16 -07:00
kshitijk4poor e513f3fb40 chore: map kshitijkapoorr@gmail.com to kshitijk4poor in contributors/emails 2026-08-26 16:06:09 +05:30
Teknium 213f46a4e3 chore: map contributor email for attribution gate 2026-08-26 03:22:37 -07:00
Deus 298ab73e7b chore: map contributor email 2026-08-26 03:22:37 -07:00
Teknium a74ffb9ee8 chore: map contributor email for projetsjsl 2026-08-26 03:21:37 -07:00
Victor Nogueira 21f34794be fix(desktop): recover remote sessions after gateway restart 2026-08-26 00:40:35 -07:00