Commit Graph

684 Commits

Author SHA1 Message Date
kshitijk4poor 06141577cd chore: AUTHOR_MAP add mrkillbob (mikedemott@Mikes-Mac-mini.local)
Contributor email for PR #93634 salvage. GitHub user mrkillbob (ID 25466867).
2026-08-28 12:38:49 +05:30
kshitijk4poor f4ae8958b5 chore: map zengzheqing noreply email for attribution (PR #96647) 2026-08-28 09:05:19 +05:30
kshitijk4poor 28ee6ac043 chore: map contributor emails for attribution (#95433, #94996) 2026-08-28 02:29:27 +05:30
Teknium 9522c4e8e2 chore: map contributor email for salvage of #95956 2026-08-27 13:48:45 -07:00
Teknium d91e4376f1 Merge pull request #65108 from NousResearch/hermes/hermes-793f4fd9
feat(skills): rewrite AgentMail optional skill CLI-first (salvages #60811)
2026-08-27 12:45:01 -07:00
kshitijk4poor f6f707b783 chore: suppress posix-gated os.kill probe in footgun scan; map rodrigogs in AUTHOR_MAP
- The stale tick-socket sweep's os.kill(pid, 0) liveness probe sits inside
  an explicit os.name == 'posix' gate (AF_UNIX nodes never exist on
  Windows) — suppress with the standard inline marker.
- contributors/emails/: rodrigo.smscom@gmail.com -> rodrigogs (author of
  the salvaged #92315 commits), unblocking check-attribution.
2026-08-27 22:06:17 +05:30
Finn763 cae58be1f5 fix(state.db): cross-backend heartbeat gates orphan sweep
A live sibling serve sharing state.db is no longer treated as a dead process by the startup orphan sweep.

Covers the sweep half of #94895. The launchd Errno 48 KeepAlive loop is not addressed here.

Credit: @Finn763
2026-08-27 11:31:59 -05:00
kshitij 59805b1dd2 chore: map contributor email for deadczarvc 2026-08-27 21:36:35 +05:30
kshitijk4poor 5dc2049bba chore: map joby@ellingtonlife.com -> gijoby (PR #91354 salvage) 2026-08-27 20:06:04 +05:30
Teknium fef1790165 chore: map contributor email for @crazyhulk 2026-08-27 07:33:36 -07:00
Teknium 2d0d2e4105 chore(contributors): map Haakam Aujla email for #60811 salvage 2026-08-27 07:11:52 -07:00
Heath Harris a71be9852e fix(kanban): close half-open tracked connection when busy_timeout PRAGMA fails
_sqlite_connect opened a connection via connect_tracked and then ran the
busy_timeout PRAGMA; if that raised, the half-open connection was abandoned
— leaking its fd AND leaving a stale entry in the sqlite_safe_read
live-connection registry (which only clears on close), permanently blocking
byte-level probes of the kanban database. Close before re-raising.

Salvaged from PR #96290 (kanban slice) with regression test.
2026-08-27 19:07:15 +05:30
kshitijk4poor 3516438832 chore(contributors): map me@kitsonkelly.com -> kitsonk 2026-08-27 17:11:56 +05:30
Teknium a588685fb9 chore(contributors): map aydinhrrs@gmail.com -> RibatTRW 2026-08-27 03:57:25 -07:00
kshitijk4poor f2d043eb08 test(cron): pin lock-first liveness + harden lock-probe failure
Follow-ups to the salvaged #95947 cron commit:

- Wrap the lock probe in its own try/except: a crashing probe is
  'unknown', not 'dead' — the pid scan still decides instead of the
  whole tri-state collapsing to None.
- Regression tests (shape adapted from #94155 by @liuhao1024): lock
  held + empty pid scan -> alive (the reported false alarm); lock
  inactive -> pid-scan fallback both ways; crashing lock probe still
  falls back.
- patch_liveness now pins the lock probe inactive by default so the
  pre-existing pid-scan tests stay deterministic on machines where a
  real gateway holds the real lock.
- contributors mapping for magnus.lundstedt@infidyne.com.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-08-27 16:10:57 +05:30
Teknium 74ec63fe18 chore: contributor email mapping for ekzhang 2026-08-27 03:24:20 -07:00
Ayush Nangia 2f0f01192d fix(cron): tree-kill script timeout descendants via agent.deadline.kill_process_tree
The script-timeout path used a site-local process-group kill, which
cannot reach a grandchild that created its OWN session (start_new_session
background jobs, watchdogs). Such descendants kept running after the job
reported failure (#71148, #59549). Migrate the timeout handler to the
unified deadline layer's kill_process_tree (#85147, d6a5cb9725): psutil
snapshots the descendant set before signalling, so own-session
grandchildren are reached too. Fallback to the site-local group kill if
the import ever fails, so the path cannot re-wedge.

The explicit script-timeout message stays the classification anchor
(#85536's contract), keeping cron timeouts distinct from provider
timeouts.

Salvage additions on review (#85125 Phase 4a):
- migrate the sibling kill site too — the cancel_event/"ownership was
  lost" path orphaned setsid grandchildren the same way (whole-bug-class
  rule); pinned by test_cancel_path_also_tree_kills
- proc.poll() early-return in _terminate_cron_script_tree so a script
  that exits right at the deadline doesn't log a spurious "no signal"
  warning (mirrors _terminate_cron_script_process); pinned by
  test_already_exited_proc_is_left_alone
- acceptance test's script timeout 1s -> 2s: interpreter startup under
  CI load could eat the whole 1s window before the spawner wrote its
  pid file
- note: kill_process_tree hard-kills (SIGKILL) immediately, whereas the
  old path gave a 1s SIGTERM grace window; intended for a deadline-
  expiry hard stop (both docstrings say "hard stop")

Based on #86791 by @ayushnangia; cherry-picked to preserve authorship.

Co-authored-by: dante32683 <dante32683@users.noreply.github.com>
Co-authored-by: supotato-ipj <supotato-ipj@users.noreply.github.com>
2026-08-27 14:28:06 +05:30
kshitij f45477b4a8 Merge pull request #86412 from kshitijk4poor/salvage/83225-overflow-clamp
fix(approval): oversized approvals.timeout crashes parallel tool batches — clamp at config read (salvage #83225/#83298, #85125 2b)
2026-08-27 14:26:38 +05:30
kshitijk4poor 5d4ad23b0d chore: map paulapsp157@gmail.com -> Aoshi-Dev (PR #96003 salvage) 2026-08-27 14:02:06 +05:30
Victor Kyriazakos db00793db7 chore(contributors): map potatosaladx@gmail.com to potatosalad (attribution for salvaged #79436/#81128 commits) 2026-08-27 15:33:32 +10:00
Teknium 26e48dd8c3 chore: map contributor email for kvnloo (#95173) 2026-08-26 21:38:20 -07:00
Teknium 89f32fe4a5 chore: map fred0m noreply email in contributors registry (#91360 salvage) 2026-08-26 17:42:50 -07:00
Casey 790e1eb6bd fix(update): pause SCM-supervised Windows gateway services before venv mutation
On Windows installs where the gateway runs as an SCM service (WinSW,
NSSM, sc.exe create), the existing pause machinery kills the gateway
process directly — and the service wrapper's failure ladder resurrects
it within seconds, re-taking the venv file locks mid-update. The update
then dies partway through dependency sync with access-denied errors.

This extends _pause_windows_gateways_for_update() to detect when a
gateway's process tree is owned by a running SCM service, and to stop
the SERVICE through sc.exe instead of killing the child:

- gateway/status.py: expose service-ownership discovery for gateway
  runtimes (find_windows_gateway_services maps validated gateway PIDs
  through process ancestry to running SCM service PIDs, with
  create-time identity checks against PID reuse).
- hermes_cli/update_cmd.py: stop verified services via sc.exe before
  venv mutation and restart them afterward. Stops wait for a stable
  SCM 'stopped' state AND for the original descendant processes to
  exit (service 'Stopped' is not proof the child released its
  handles). Failure to prove ownership, stop a service, or restart it
  fails closed; rollback restores attempted services, and rollback
  failures are surfaced rather than swallowed.
- Fail-closed throughout: unreadable identities, ambiguous ancestry,
  or a service that will not reach a stable state abort the update
  before any file mutation.

Complements #37039 (gateway-only concurrent instances no longer abort):
that fix lets the update proceed past the gate; this one makes the
pause actually stick when the gateway is service-supervised.

Note: tests/gateway/test_status.py::TestReadProcessCmdlinePsFallback::
test_ps_fallback_when_proc_unavailable fails on Windows on current main
before this change as well (POSIX ps fallback asserted on a platform
without it); all other touched suites pass (155 passed, 5 skipped).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 16:45:31 -07:00
Finn763 de34c746cc chore: map contributor emails 2026-08-26 15:51:22 -07:00
AlexGabbia b3e477f304 fix(update): wait for resumed Windows gateway before failing fleet check
The post-update fleet version check slept 2s and probed once. On Windows the
resume path relaunches the gateway detached, and it needs ~10s to boot (the
Telegram polling reconnect) before it stamps gateway_state.json or answers the
control socket. That race reported "no rows" for a healthy resume, exited 1,
and triggered a full retry that re-killed the gateway the first attempt had
just started — leaving it down and surfacing "Update failed (exit 1)".

Poll a bounded window (up to 30s) for the resumed gateway to publish its
identity, and only treat a persistently empty snapshot as verification
failure. The fail-closed contract from #93406 is preserved: a gateway that
genuinely never comes back still exits 1.
2026-08-26 15:05:32 -07:00
Teknium 3a4c309cfb chore: map contributor email for salvage of #88470 2026-08-26 10:10:11 -07:00
Teknium 4e8e5f84bc chore: contributor mappings for #95633/#95652 salvage 2026-08-26 09:51:55 -07:00
etzelvon dbbed456c1 style(desktop): satisfy curly + padding-line lint for guarded switch
- braces for single-statement if(confirmNotificationId) guards
- blank line before return in staleness branch (padding-line-between-statements)
2026-08-26 08:39:45 -07:00
Teknium f4e8eb1568 chore: map A-Snegin attribution 2026-08-26 07:28:10 -07:00
Teknium 1145fcaed6 chore: map contributor email for deepeet-git 2026-08-26 06:26:02 -07:00
Teknium 279726cc2f chore: map contributor email for wiseconnex 2026-08-26 04:51:41 -07:00
Teknium 6bbae974d5 chore: map justinjohnson25600 and BrunoBza contributor emails 2026-08-26 04:49:22 -07:00
Teknium 20d33e385a fix(web_server): a dashboard started without a build recovers the moment one appears (#82614)
mount_spa's WEB_DIST.exists() check ran ONCE at mount time: a long-lived
'hermes dashboard --skip-build' that survived a git pull (or launched
before the first build) installed a permanent no_frontend catch-all and
answered 404 'Frontend not built' on every route forever — even after
npm run build completed. Remote Desktop clients saw ERR_EMPTY_RESPONSE.

The missing-dist branch is now reserved for the headless-serve contract
only. The SPA routes mount unconditionally and already cope with a
missing dist per-request (_serve_index returns the same 404 JSON when
index.html is unreadable; the /assets mount gains check_dir=False so
StaticFiles 404s instead of raising at mount). The dashboard recovers
the moment a build lands on disk — no restart needed.

Direction from #82666 by @codexbt (his PR's rebase dropped the product
hunk, leaving only the test; the test is cherry-picked as-is and this
commit restores the behavior it pins, adapted to the current mount_spa
shape: headless guard preserved, per-request recovery instead of a
per-request exists() check).
2026-08-26 04:48:58 -07:00
Teknium 7a10d91b29 chore: map contributor email for notkisk 2026-08-26 04:14:16 -07:00
kshitijk4poor e513f3fb40 chore: map kshitijkapoorr@gmail.com to kshitijk4poor in contributors/emails 2026-08-26 16:06:09 +05:30
Teknium 213f46a4e3 chore: map contributor email for attribution gate 2026-08-26 03:22:37 -07:00
Deus 298ab73e7b chore: map contributor email 2026-08-26 03:22:37 -07:00
Teknium a74ffb9ee8 chore: map contributor email for projetsjsl 2026-08-26 03:21:37 -07:00
Victor Nogueira 21f34794be fix(desktop): recover remote sessions after gateway restart 2026-08-26 00:40:35 -07:00
Jaime Marques 14d16c2578 fix(desktop): clarify primary SSH reuse failures 2026-08-26 00:40:35 -07:00
Teknium 4fc4da2531 chore: map contributor email for salvaged PR #88544 2026-08-26 00:40:35 -07:00
Ravi Tharuma 949f5169de fix(desktop): treat ticket 401 as sign-in when native tokens are unreadable 2026-08-26 00:40:35 -07:00
kshitijk4poor 4ba2608524 fix(compressor): widen empty-content abort to sibling no-response shapes + snapshot state field
Follow-up to PR #94531 salvage:
- classify the auxiliary boundary's terminal 'None response' /
  'invalid response' errors (#7264) into the same empty-content abort
  carve-out so those shapes also preserve the session (#94459's wider
  classification, sibling shapes from #94448)
- register _last_summary_empty_content_failure in
  _COMPRESSOR_ATTEMPT_STATE_FIELDS so pre-commit hard-cancel rollback
  restores the flag (conversation_compression snapshot allow-list)
- tests: cooldown re-entry keeps aborting; both sibling shapes abort
- attribution: map zhangyswx@163.com -> YusenZhang0601
2026-08-26 13:02:27 +05:30
Tom 2ed39365d6 fix(desktop): thread eager profile metadata through registry enumeration
Never-interacted remote bots painted as bare handles because roster rows
carried only profile names: display_name/title/ui_meta/has_avatar were
fetched lazily on first interaction (#91365). Thread credential-free
profile metadata from the enumeration-time /api/profiles body through
enumerateRegistryAgentSources (main.ts) and buildAgentRoster
(connection-registry.ts), keeping it attached to the connection-qualified
row across the same-install collapse. The plugin.js botRosterMeta half of
the original PR is dropped — superseded by landed #92731.

Fixes #91365
Salvaged (partial) from #92708.
2026-08-26 00:22:27 -07:00
Tilly-YL 2952119bce fix(desktop): remember selected profile across restarts
The profile rail's live workspace switch never persisted the selection,
so the Desktop always booted back into the previous startup profile
(#79886). Route the successful primary-backend activation through a new
persistence-only hermes:profile:remember IPC (validated
writeActiveDesktopProfile) that records the choice WITHOUT tearing down
the backend or reloading the window like hermes:profile:set does.
Registry-source picks name another source's profiles and do not touch
the startup preference. Reapplied semantically over three weeks of
main.ts/preload.ts drift (selectProfile now routes through
activateOnCurrentSource, #91349/#91365 seams).

Fixes #79886
Salvaged from #79888.
2026-08-26 00:22:27 -07:00
Teknium e0210ab6c9 chore: contributor email mappings for salvage class-4 2026-08-26 00:21:49 -07:00
LovePlayCode 1808d33a24 fix(desktop): keep owner-routed tile gateways out of idle prune
Bot chats stay on a secondary while chrome stays on the launch profile.
The keep-set only counted busy sessions, so idle prune closed the tile
socket and resume spun forever. Keep open tiles, route catalog reads to
the owner, and hydrate model/provider from resume.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-26 00:21:29 -07:00
Teknium 708b24513d chore: map contributor email for jwtor7 2026-08-25 23:33:46 -07:00
Teknium a7e1410faf chore: map Kavyrocom attribution 2026-08-25 23:15:49 -07:00
Teknium 2664599644 test(gateway): prove the profile_route_rejected sentinel is observable
Review follow-up for the salvaged #94848: the ProfileRouteRejected marker
looked write-only inside the primary handler. Document that
_handle_message's ingress gate reads the same marker to drop the message
fail-closed, and add a regression test showing the rejected route is
stamped once, dispatch falls back to the default home, and routing is not
re-run on redelivery.
2026-08-25 23:15:15 -07:00